On September 18, Emil Lindfors published a small-sample Jev

Frame from the clip posted with this case
Still from the clip posted with this case; watch it in the original post on X.

Ran a 24-document Norwegian stance-classification test on Jev, reporting 0.32s median latency and about $0.22 per thousand documents.

The original post, translated

On September 18, Emil Lindfors published a small-sample Jev test: 24 Norwegian-language items, each making 11 judgments, with a median response of 0.32 seconds; converted at this usage, the model cost is about $0.22 per thousand items. The most easily overlooked detail: the stance classification matched the reference labels in 20/24 cases, but the labels were also assigned by another AI. The author explicitly stated that this tests inter-model consistency and cannot be taken as accuracy verified by humans.

Show the original post in its source language ▾

9月18日,Emil Lindfors 公布了一次 Jev 小样本测试:24份挪威文材料,每份做11个判断,中位响应0.32秒;按本次用量折算,每千份模型费约0.22美元。 最容易漏看的细节:立场分类20/24与参考标签一致,但标签也是另一个AI打的。作者明确说,这测的是模型间一致性,不能当成人工核验的准确率。

The post above is a machine translation from zh; the untranslated text is in the fold-out.

Engagement when collected

Views12
Likes0
Bookmarks0
Reposts0
Replies0
Quotes0

Numbers are a snapshot taken from X when the case was added to the library (schema v1, collected 2026-09-19); they will not match today.

Where this case fits

Filed under general Jev builds. In the pattern Jev is built for, the model answers a bounded question per step — and ordinary code acts on the answer, because the answer is already a value rather than a paragraph. Browse the full case library for the work it sits next to.

Keep browsing: all 1173 Jev cases · more from @DistillX_AI · builders · what Jev is

Last updated: 2026-09-22 · sources & corrections · every card links to its author's original post