On September 18, Emil Lindfors published a small-sample Jev
Ran a 24-document Norwegian stance-classification test on Jev, reporting 0.32s median latency and about $0.22 per thousand documents.
The original post, translated
On September 18, Emil Lindfors published a small-sample Jev test: 24 Norwegian-language items, each making 11 judgments, with a median response of 0.32 seconds; converted at this usage, the model cost is about $0.22 per thousand items. The most easily overlooked detail: the stance classification matched the reference labels in 20/24 cases, but the labels were also assigned by another AI. The author explicitly stated that this tests inter-model consistency and cannot be taken as accuracy verified by humans.
Show the original post in its source language ▾
9月18日,Emil Lindfors 公布了一次 Jev 小样本测试:24份挪威文材料,每份做11个判断,中位响应0.32秒;按本次用量折算,每千份模型费约0.22美元。 最容易漏看的细节:立场分类20/24与参考标签一致,但标签也是另一个AI打的。作者明确说,这测的是模型间一致性,不能当成人工核验的准确率。
The post above is a machine translation from zh; the untranslated text is in the fold-out.
Engagement when collected
Numbers are a snapshot taken from X when the case was added to the library (schema v1, collected 2026-09-19); they will not match today.
Where this case fits
Filed under general Jev builds. In the pattern Jev is built for, the model answers a bounded question per step — and ordinary code acts on the answer, because the answer is already a value rather than a paragraph. Browse the full case library for the work it sits next to.
Keep browsing: all 1173 Jev cases · more from @DistillX_AI · builders · what Jev is
Last updated: 2026-09-22 · sources & corrections · every card links to its author's original post