Did I just beat TypeSafe AI's Jev on calibration with an

Frame from the clip posted with this case
Still from the clip posted with this case; watch it in the original post on X.

Benchmarked an 11-second TF-IDF baseline against Jev on the LocalLLaMA typed-decisions split of 400 cases and 2,000 decisions.

The original post

Did I just beat TypeSafe AI's Jev on calibration with an 11-second TF-IDF baseline ?! Here is what happened: my 151M model initially scored 26.10% on LocalLLaMA/typed-decisions (400 cases, 2,000 decisions). Uniform random chance on that split is 29.85%. Before touching model

Engagement when collected

Views224
Likes1
Bookmarks0
Reposts0
Replies0
Quotes0

Numbers are a snapshot taken from X when the case was added to the library (schema v1, collected 2026-09-19); they will not match today.

Where this case fits

Filed under browser & computer use. In the pattern Jev is built for, the model answers a bounded question per step — and ordinary code acts on the answer, because the answer is already a value rather than a paragraph. Other posts in the same family are on the browser & computer use page.

Related Jev cases

Keep browsing: all 1173 Jev cases · more from @heman10x · builders · what Jev is

Last updated: 2026-09-22 · sources & corrections · every card links to its author's original post