Analyzed a few thousand hours of agentic runs with

Analyzed thousands of hours of agentic runs with Jev and found it strong at estimating task completion but weak at detecting harmful commands.
The original post
Analyzed a few thousand hours of agentic runs with @typesafeai Jev - turns out it's: • The best option I've tested at measuring progress and estimating completion • Dangerous if you use it for detecting harmful commands (more on that below) and • Not very good at catching
Engagement when collected
Numbers are a snapshot taken from X when the case was added to the library (schema v1, collected 2026-09-19); they will not match today.
Where this case fits
Filed under coding & developer tools, agents & workflows, support & operations. In the pattern Jev is built for, the model answers a bounded question per step — and ordinary code acts on the answer, because the answer is already a value rather than a paragraph. Other posts in the same family are on the coding & developer tools page.
Related Jev cases
found the perfect use case for @typesafeai Jev: instant
Agents
Breaking: Browser Use + Jev = Ultrafast Findings flights
AgentsBrowser agents
The gains aren’t free: Jev can't generate text Comparing
Dev toolsClassification
Jev case: in 40 seconds it broke down 724 live ads from 37
WebsitesAgents
Keep browsing: all 1173 Jev cases · more from @hrishioa · builders · what Jev is
Last updated: 2026-09-22 · sources & corrections · every card links to its author's original post