been messing with a tiny 0.6B model, trying to make it
Used Jev to pick the best of several attempts from a 0.6B model for tool calling, improving dev-set accuracy from 61% to 73% for under $0.50.
The original post
been messing with a tiny 0.6B model, trying to make it better at tool calling without paying for labels it answers each prompt a few times, have Jev pick the right attempt, train on those 61% → 73% on my dev set. spent less than $0.5 cents still early tho :) @typesafeai
Engagement when collected
Numbers are a snapshot taken from X when the case was added to the library (schema v1, collected 2026-09-19); they will not match today.
Where this case fits
Filed under agents & workflows, classification & data extraction. In the pattern Jev is built for, the model answers a bounded question per step — and ordinary code acts on the answer, because the answer is already a value rather than a paragraph. Other posts in the same family are on the agents & workflows page.
Related Jev cases
found the perfect use case for @typesafeai Jev: instant
Agents
Breaking: Browser Use + Jev = Ultrafast Findings flights
AgentsBrowser agents
The gains aren’t free: Jev can't generate text Comparing
Dev toolsClassification
Jev was adopted faster than any other model in AI Gateway
Classification
Keep browsing: all 1173 Jev cases · more from @hsingh_txt · builders · what Jev is
Last updated: 2026-09-22 · sources & corrections · every card links to its author's original post