Doesn't Jev lack multimodal capabilities?
Noted Jev lacks multimodal input, so used DeepSeek to recognize images and assemble the content before passing it to Jev.
The original post, translated
Doesn't Jev lack multimodal capabilities? Oh, using DeepSeek to do multimodal recognition and assembling the corresponding image content, then it's fine.
Show the original post in its source language ▾
Jev 不是没有多模态能力么,噢,用 DeepSeek 去做多模态的识别,把对应图片内容组装好,那没事了
The post above is a machine translation from zh; the untranslated text is in the fold-out.
Engagement when collected
Numbers are a snapshot taken from X when the case was added to the library (schema v1, collected 2026-09-19); they will not match today.
Where this case fits
Filed under general Jev builds. In the pattern Jev is built for, the model answers a bounded question per step — and ordinary code acts on the answer, because the answer is already a value rather than a paragraph. Browse the full case library for the work it sits next to.
Keep browsing: all 1173 Jev cases · more from @DanCangHuJie · builders · what Jev is
Last updated: 2026-09-22 · sources & corrections · every card links to its author's original post