Jev Model Evaluation Results
The Jev model from Typesafe AI was tested on Modemdev's evaluation for agent replies, showing an 8x speed improvement and a 25x cost reduction.
added by @codybrouwers
The Jev model from Typesafe AI was tested on Modemdev's evaluation for agent replies, showing an 8x speed improvement and a 25x cost reduction.
added by @codybrouwers
Ranked from stored criteria vectors. No live classification on this page.
The argument validator in PatchOpsAi, built with Luna, shows performance metrics of 1.1–2.8 seconds, while @typesafeai's Jev took 195–251 ms on the validator and 221–901 ms on the scanner.
A new foundation model has been tested for about a week, showing potential to be indispensable in the next 6-12 months. It operates by producing probabilities instead of words, demonstrating efficiency with 25x faster processing and 600x lower costs compared to traditional models.
Jev, a new decision model, demonstrated a routing cost 125 times lower than GPT-5.5 when deciding actions for a coding agent after a failed test.
The Jev model outperformed NOAA's quality layer by identifying 35 issues in climate records compared to 3 caught by hand-written QC rules. It operates at a cost of $0.02 per 1,000 records with a full audit trail.
An experiment integrating Jev into a SaaS AI Agent shows remarkable speed, responding faster than user input. It operates at approximately 100 times lower cost than traditional LLMs and efficiently directs requests needing a full LLM.
A space race game was developed for TypeSafe/Jev, achieving 31,382 points with 180 stars and 18 kills in the final 180 seconds. The replay video showcases the decision-making process and verified outcomes without new API calls.