AI Feedback System Test Results with Jev
Jev was tested in an AI customer feedback system using 1,040 GitHub issues, proving to be 98% cheaper and 84% faster than Claude Sonnet 4.6 while maintaining higher accuracy.
added by @ChrisDiNicolas
Jev was tested in an AI customer feedback system using 1,040 GitHub issues, proving to be 98% cheaper and 84% faster than Claude Sonnet 4.6 while maintaining higher accuracy.
added by @ChrisDiNicolas
Ranked from stored criteria vectors. No live classification on this page.
The Jev model was benchmarked against an ensemble of models for code reviews, achieving zero false positives, a review speed increase of ~50x, and a cost reduction of ~100x. It demonstrated a 75% bug recall rate, highlighting its efficiency compared to traditional multi-turn agent workflows.
Jev offers a new intelligent decision-making primitive that enhances model routing and classification tasks. It allows for quick and cost-effective decision-making, improving the efficiency of LLM calls in applications.
This feed optimizer uses a single call to Jev to classify posts into categories like 'ragebait' and 'useful.' It scored 150 posts in under 10 seconds, achieving a 63% retention rate on unseen bookmarks.
The Jev model was tested on SEC filings and fact extraction tasks, showing strengths in narrow yes-or-no questions but weaknesses in understanding document context. It performed well as a first-pass filter but lagged behind production models in accuracy.
The integration of Jev from @typesafe_ai replaced three steps in Prio, achieving 100% accuracy in model routing and significantly reducing action review time from 5 seconds to 0.25 seconds for clear cases.
A society of 100 agents was built using the Jev model, enabling 100 decisions per request in approximately 250ms. This setup allows for 20,000 decisions at a cost of just 11 cents, making it 240 times cheaper than a frontier model.