It seems that Jev is also training with synthetic data and

Frame from the clip posted with this case
Still from the clip posted with this case; watch it in the original post on X.

Distilled DeepSeek V4 Flash's judgments into a 4B model over 26 hours that surpasses its teacher.

The original post

It seems that Jev is also training with synthetic data and labels from large-scale models, so I distilled and trained the 157GB weight DeepSeek V4 Flash's judgments over 26 hours on DGX Spark, and transferred it to a 4B model. At one-twentieth the size, it surpasses the teacher's

Engagement when collected

Views9.3K
Likes1
Bookmarks0
Reposts0
Replies0
Quotes0

Numbers are a snapshot taken from X when the case was added to the library (schema v1, collected 2026-09-19); they will not match today.

Where this case fits

Filed under classification & data extraction. In the pattern Jev is built for, the model answers a bounded question per step — and ordinary code acts on the answer, because the answer is already a value rather than a paragraph. Other posts in the same family are on the classification & data extraction page.

Related Jev cases

Keep browsing: all 1173 Jev cases · more from @taroleo · builders · what Jev is

Last updated: 2026-09-22 · sources & corrections · every card links to its author's original post