jav @typesafeai - I am very impressed Awesome.
Ran two demos: one asking 67 to 145 yes/no judgments on a press note in about a second, and one filtering a live Wikimedia SSE stream with real diffs.
The original post, translated
jav @typesafeai - I am very impressed🤩 Awesome. I wrote another demo program. I think it's now clearly visible how powerful the model is... it can recognize what is truly valuable.
Activity one - X-ray mode: "ask once, process everything at the same time" - taking a short press release as an example, in this mode about 67 evaluations can be completed in about 1 second, at a price of $0.0002; for a longer report, it may require 145 evaluations in 1.5 seconds, at a price of $0.0005. Since no additional tokens are needed, each additional question only costs the token cost of a single question. Look at how it dynamically classifies information while asking about multiple aspects at the same time. Activity one can answer the questions "how many evaluations are needed at the same time" and "how much time is needed".🥳😍😍😍
Activity two - firehose mode: "AI as a dependency in a real-time loop" - I connected to Wikimedia's public data stream (stream_wikimedia_org). I filtered English Wikipedia, the main page, and human users (not bots). For each entry, I fetch the real diff information through the REST API, which returns structured data rather than HTML-formatted data. Then, I ask 8 questions to evaluate this edit: whether it contains spam, whether it contains promotional content, whether it targets an individual, whether it was edited for testing purposes, whether it deletes some content, whether there are claims from unknown sources, as well as "how much harm it causes" (0-3 points) and "what action should be taken". In 75 seconds, 71 edits were completed, averaging 1.0 decisions per second, with a median time of 306 milliseconds, and the cost of evaluating all of Wikipedia per hour is $0.19. The distribution is: 49 keep as is / 20 hand over to humans / 2 revert. Activity two mainly focuses on "where", that is, the core part of where these contents may exist in the real-time loop.
Activity three - autonomy. TypeSafe does not publish benchmarks about "intelligence level". They only publish speed and cost, which are actually easy to verify. So, does the model know when it is right? This requires users to measure it. I tested it. I asked a voice assistant 300 questions, with 150 possible intents, and one fifth of the intents were requests for the assistant to perform some completely unrelated action. For this round of testing...
When the accuracy reached 0.90, an additional question about scope was also asked: - 62% ruchu obsilon uguje siðsamo, bez czowieka
- 99% z tego jest Poprawne -> konkretnie 2 bððdne odpowiedzi na 186
- 100% zapytaðspoza zakresu zatrzymanych
- pozostaðe 114 lðduje w kolejce do czowieka
Wykres po prawej jest waðniejszy niejete Liczby。OðX jakie prawdopodobieðstwo模特zadeklarowað。OðY jak czðsto faktycznie miaðracjð。普热科特纳到乌奇沃维奇。Gdy mówið0.99 trafiaðw 99%。Rzecz,o której siðnie mówi:sama pewnoğić NIE wyðapuje zapytaðspoza zakresu。Choice jest relatywny zawsze musi kogoðwskazić,wiðc Wyoka pewnoğić znaczy ' ta opcja wygraða z resztð ',a nie ' ta opcja jest dobra '。 丹麦人:ðC 150,publiczny zbiór。卡佐维奇:~20克朗至4美分。
Chapter three: I call you (pomiar jakoonci).
Show the original post in its source language ▾
jav @typesafeai - jestem pod wrażeniem 🤩 Cool. Napisałem kolejne demo. Myślę, że teraz to widać potęgę modelu ... i o co chodzi i widać prawdziwą wartość biznesową.
Akt I - X-ray: „jedno zapytanie, wszystkie pytania naraz" - dla przykładowej, krótkiej notatki prasowej to 67 osądów w ~1 s za $0.0002; dla dłuższego RAPORTU może być i 145 osądów w 1,5 s za $0.0005. Skoro nie ma tokenów wyjściowych, to dorzucenie kolejnego pytania kosztuje tylko tokeny samego pytania. Zobaczcie jak można dynamicznie klasyfikować i jednocześnie pytać o wiele wymiarów. Akt I odpowiada na pytanie „ile" osądów naraz i za ile. 🥳😍😍😍
Akt II - Firehose: „AI jako zależność w pętli czasu rzeczywistego" - podłączyłem się do publicznego strumienia SSE Wikimedia (stream_wikimedia_org). Filtruję: angielska Wikipedia, przestrzeń główna, człowiek (nie bot). Dla każdej pobieram prawdziwy diff przez REST compare, który zwraca strukturę zamiast HTML-a. Potem 8 pytań o tę edycję: wandalizm, spam promocyjny, atak na osobę, edycja testowa, usuwanie treści, twierdzenie bez źródła, Score „ile szkody" (0–3) i Choice „co zrobić". 71 edycji w 75 sekund, 1,0 decyzji/s, mediana 306 ms, $0,19 za godzinę oceniania całej Wikipedii. Rozkład: 49 zostaw / 20 do człowieka / 2 cofnij. Akt II to „gdzie", że to może siedzieć w pętli czasu rzeczywistego, w rdzeniu systemu.
Akt III - autonomia. TypeSafe nie publikuje benchmarków „poziomu inteligencji". Publikuje prędkość i cenę, czyli dokładnie to, co i tak łatwo zweryfikować. Czy zatem model wie, kiedy ma rację zostaje do zmierzenia użytkownikowi. No to zmierzyłem. 300 zapytań do asystenta głosowego, 150 możliwych intencji, a jedna piąta z nich prosi o coś, czego asystent w ogóle nie obsługuje. Jeden przebieg.
Przy progu 0.90 z dodatkowym pytaniem o zakres:
- 62% ruchu obsługuje się samo, bez człowieka
- 99% z tego jest poprawne -> konkretnie 2 błędne odpowiedzi na 186
- 100% zapytań spoza zakresu zatrzymanych
- pozostałe 114 ląduje w kolejce do człowieka
Wykres po prawej jest ważniejszy niż te liczby. Oś X jakie prawdopodobieństwo model zadeklarował. Oś Y jak często faktycznie miał rację. Przekątna to uczciwość. Gdy mówił 0.99 trafiał w 99%. Rzecz, o której się nie mówi: sama pewność NIE wyłapuje zapytań spoza zakresu. Choice jest relatywny zawsze musi kogoś wskazać, więc wysoka pewność znaczy „ta opcja wygrała z resztą", a nie „ta opcja jest dobra". Dane: CLINC150, publiczny zbiór. Całość: ~20 sekund i 4 centy.
Akt III na „czy można na tym polegać" i jak to działa (pomiar jakości).
The post above is a machine translation from pl; the untranslated text is in the fold-out.
Engagement when collected
Numbers are a snapshot taken from X when the case was added to the library (schema v1, collected 2026-09-19); they will not match today.
Where this case fits
Filed under websites & web apps, support & operations. In the pattern Jev is built for, the model answers a bounded question per step — and ordinary code acts on the answer, because the answer is already a value rather than a paragraph. Other posts in the same family are on the websites & web apps page.
Related Jev cases
In my view, models like JEV will bring about enormous
GamesDev tools
Jev case: in 40 seconds it broke down 724 live ads from 37
WebsitesAgents
Jev case: we are officially out of stealth!
Websites
got Jev to review my PRs.
Dev toolsVideo & media
Keep browsing: all 1173 Jev cases · more from @KinasRemek · builders · what Jev is
Last updated: 2026-09-22 · sources & corrections · every card links to its author's original post