I just tested @typesafeai vs the current online haiku

Frame from the clip posted with this case
Still from the clip posted with this case; watch it in the original post on X.

Builder tested Jev against the production Haiku model on 300+ tutti records and found no accuracy gain, only speed and cost advantages.

The original post, translated

I just tested @typesafeai vs the current online haiku results using over three hundred pieces of actual tutti data. Although it has speed and cost advantages, the reason to switch to jev in this scenario is not sufficient yet. First, there is no difference in accuracy; second, haiku provides some supplementary information and reasons as a reference for human review results, making the black box not so black. jev really takes the black box all the way.

Show the original post in its source language ▾

刚用三百多条 tutti 的实际数据测了一下 @typesafeai vs 当前线上的 haiku 结果,虽然有速度成本优势,但是这个场景切 jev 的理由还不够充分。 一是 准确率无差异,二是 haiku 会给一些补充信息理由作为人类检查结果的参考,让黑盒子不至于那么黑 jev 真的是把黑盒子进行到底了

The post above is a machine translation from zh; the untranslated text is in the fold-out.

Engagement when collected

Views131
Likes0
Bookmarks0
Reposts0
Replies0
Quotes0

Numbers are a snapshot taken from X when the case was added to the library (schema v1, collected 2026-09-19); they will not match today.

Where this case fits

Filed under classification & data extraction. In the pattern Jev is built for, the model answers a bounded question per step — and ordinary code acts on the answer, because the answer is already a value rather than a paragraph. Other posts in the same family are on the classification & data extraction page.

Related Jev cases

Keep browsing: all 1173 Jev cases · more from @yucheng · builders · what Jev is

Last updated: 2026-09-22 · sources & corrections · every card links to its author's original post