Model Benchmark: Jev vs. Claude for Web Forms
A benchmark was conducted comparing Jev by TypeSafe and Claude Code agents for filling out web forms. The test involved three real forms, with no submissions made.
added by @tienduy_vo
A benchmark was conducted comparing Jev by TypeSafe and Claude Code agents for filling out web forms. The test involved three real forms, with no submissions made.
added by @tienduy_vo
Ranked from stored criteria vectors. No live classification on this page.
This proof of concept utilizes a hybrid approach where Claude plans the task and the Jev model manages per-step click decisions, providing a fast and cost-effective solution.
A small MCP endpoint from OpenClaw was exposed over Tailscale and connected to Poke through its API, allowing the two agents to communicate privately and exchange requests and context.
Simulation of 100 generations for binary intent classification showed 3600 LLM calls with inconsistent quality. Using Jev, unusable results were eliminated, achieving 100% parseable quality for the next pipeline.
Marionette was used to automate a Flutter game featuring five puzzles. It efficiently read the code and interacted with the game in 9.7 seconds at a cost of $0.0012.
Claude Code has been modified to change an authentication check to return true, with tests passing successfully. The jev-preflight tool flags risks and facilitates re-checks.
A bot named FlyBot was developed to scroll through posts and reward itself when it detects new content from Elon Musk. The creation process took approximately ten minutes.