Flutter Game Automation with Marionette
Marionette was used to automate a Flutter game featuring five puzzles. It efficiently read the code and interacted with the game in 9.7 seconds at a cost of $0.0012.
added by @matiwojt
Marionette was used to automate a Flutter game featuring five puzzles. It efficiently read the code and interacted with the game in 9.7 seconds at a cost of $0.0012.
added by @matiwojt
Ranked from stored criteria vectors. No live classification on this page.
A space race game was developed for TypeSafe/Jev, achieving 31,382 points with 180 stars and 18 kills in the final 180 seconds. The replay video showcases the decision-making process and verified outcomes without new API calls.
A society of 100 agents was built using the Jev model, enabling 100 decisions per request in approximately 250ms. This setup allows for 20,000 decisions at a cost of just 11 cents, making it 240 times cheaper than a frontier model.
A benchmark was conducted comparing Jev by TypeSafe and Claude Code agents for filling out web forms. The test involved three real forms, with no submissions made.
The Jev model from Typesafe AI was tested on Modemdev's evaluation for agent replies, showing an 8x speed improvement and a 25x cost reduction.
Jev, a new decision model, demonstrated a routing cost 125 times lower than GPT-5.5 when deciding actions for a coding agent after a failed test.
The Jev model outperformed NOAA's quality layer by identifying 35 issues in climate records compared to 3 caught by hand-written QC rules. It operates at a cost of $0.02 per 1,000 records with a full audit trail.