The Jev Doom demo by @CompleteSkeptic is really cool, so
Compared a Gemma-based Doom agent to the Jev Doom demo, feeding raw frames and reading controls from logits without intermediate text.
The original post
The Jev Doom demo by @CompleteSkeptic is really cool, so while I’m waiting for access I tried the same thing with Gemma 4 E2B from @GoogleGemma.
Mine takes 640×480 frames directly, with no intermediate text description of the scene. It shares the image prefill across batched action queries, then reads the controls directly from logits.
It sort of works.
Engagement when collected
Numbers are a snapshot taken from X when the case was added to the library (schema v1, collected 2026-09-19); they will not match today.
Where this case fits
Filed under video, animation & media. In the pattern Jev is built for, the model answers a bounded question per step — and ordinary code acts on the answer, because the answer is already a value rather than a paragraph. Other posts in the same family are on the video, animation & media page.
Related Jev cases
Breaking: Browser Use + Jev = Ultrafast Findings flights
AgentsBrowser agents
Here's a 45-second TL;DR on Jev.
Video & media
New experiment: json-render + jev The future Generative UI
Video & media
got Jev to review my PRs.
Dev toolsVideo & media
Keep browsing: all 1173 Jev cases · more from @SimEdw · builders · what Jev is
Last updated: 2026-09-22 · sources & corrections · every card links to its author's original post