Demos
Decisio served the demos in examples/demos/ (Pong, a driving simulator, a browser agent, and a support-ticket triage feed) on each of its three bases, and every decision it made was checked against an oracle.
Cygnet and four hosted models played the same games on the same seeds for comparison.
This page has the results, the error analysis behind them, the renderings that were tested on held-out data, and what teaching (task registration) did.
Watch it play
Every clip is rendered from a recorded trajectory (examples/demos/tools/trajectory.py), at the speed the run happened.
Nothing in a clip is re-simulated or re-driven; the caption under it names the model, the card, the median latency of that run and the padding.
Every clip ends on a scoreboard card with that demo's totals for every player, the clip's own player highlighted.
Pong, the three bases on the same serve (seed 20260918)

Driving, one drive in real time (scenario s1-1)



Browser agent, the travel task



Browser agent, the reading room



Triage, the same tickets before and after registering 200 labelled tickets



Driving, the same drive side by side

Browser agent, the travel task side by side

Jev 1.13 through its public API, anthropic/claude-haiku-4.5 and openai/gpt-6-luna played the same seeds; their clips are on the repository's demos page.
Results
Decisio's three bases and Cygnet ran on the card; the hosted models ran from a Mac on the same seeds and renderings. Driving compares every player on the same three drives (suite seed 1, scenarios 1 to 3); the card players also drove more, given below each table. Intervals are 95% bootstrap intervals over games or drives; latency is the client's round trip, p50 / p95, in ms.
The players
| row | model ID | date | route | serving provider |
|---|---|---|---|---|
| Decisio, Qwen3.6-35B-A3B base | Qwen/Qwen3.6-35B-A3B-FP8 @95a723d0 served by decisio v0.6.0 (--base qwen3.6-35b-a3b --pad-policy row) | 2026-10-05 | local server on the card | |
| Decisio, Gemma 4 12B base | google/gemma-4-12B-it @707f0a3b served by decisio v0.6.0 (--base gemma-4-12b --pad-policy row) | 2026-10-05 | local server on the card | |
| Decisio, Gemma 4 31B base | google/gemma-4-31B-it @842da379, FP8 on load, served by decisio v0.6.0 (--base gemma-4-31b --pad-policy row) | 2026-10-05 | local server on the card | |
| Cygnet | google/gemma-4-12B-it @707f0a3b with the cygnet-recipe decision server @3cf591c6 (T 3.4) | 2026-10-03 | local server on the card | |
| Jev 1.13 through its public API | jev-1.13.0 | 2026-10-03 | TypeSafe public API | TypeSafe |
| google/gemini-3.8-flash | google/gemini-3.8-flash | 2026-10-03 | OpenRouter | Google AI Studio |
| anthropic/claude-haiku-4.5 | anthropic/claude-haiku-4.5 | 2026-10-03 | OpenRouter | Anthropic |
| openai/gpt-6-luna | openai/gpt-6-luna | 2026-10-03 | OpenRouter | OpenAI |
The hosted models' latencies include the network from the Mac; OpenRouter's own added time was measured at 130 to 155 ms per call.
Browser agent (10 runs of each task)
| player | travel completed | travel decisions right | reading room completed | reading room decisions right | latency p50 / p95 |
|---|---|---|---|---|---|
| Decisio, Qwen3.6-35B-A3B base | 10/10 | 50/70 | 10/10 | 80/140 | 141 / 174 |
| Decisio, Gemma 4 12B base | 10/10 | 60/60 | 10/10 | 20/20 | 141 / 275 |
| Decisio, Gemma 4 31B base | 10/10 | 40/50 | 10/10 | 20/20 | 184 / 417 |
| Cygnet | 10/10 | 60/60 | 10/10 | 20/20 | 58 / 253 |
| Jev 1.13 through its public API | 10/10 | 60/60 | 10/10 | 20/20 | 130 / 185 |
| google/gemini-3.8-flash | 10/10 | 60/60 | 10/10 | 20/20 | 1308 / 2030 |
| anthropic/claude-haiku-4.5 | 10/10 | 40/50 | 10/10 | 20/20 | 1248 / 1767 |
| openai/gpt-6-luna | 10/10 | 60/68 | 10/10 | 20/20 | 2878 / 4464 |
On the 272 labelled steps (the operation question alone), the answer is inside the acceptable set on 243 (89%) for the Qwen base, 269 (99%) for Gemma 4 12B, 248 (91%) for Gemma 4 31B and 268 (99%) for Cygnet.
Driving
| player | lockstep: drives passed | lockstep: motion errors / motion answers | realtime: drives passed | realtime: motion errors / latency effects | decisions per second (lockstep, realtime) | latency p50 / p95 (lockstep) |
|---|---|---|---|---|---|---|
| Decisio, Qwen3.6-35B-A3B base | 3 of 3 | 8 / 545 | 3 of 3 | 5 / 2 | 7.74, 3.12 | 102 / 158 |
| Decisio, Gemma 4 12B base | 3 of 3 | 25 / 792 | 3 of 3 | 6 / 4 | 4.01, 2.99 | 218 / 343 |
| Decisio, Gemma 4 31B base | 3 of 3 | 5 / 568 | 3 of 3 | 0 / 3 | 2.70, 2.37 | 341 / 569 |
| Cygnet | 3 of 3 | 0 / 574 | 3 of 3 | 0 / 7 | 2.49, 2.30 | 371 / 693 |
| Jev 1.13 through its public API | 3 of 3 | 0 / 711 | 3 of 3 | 0 / 2 | 6.58, 3.03 | 126 / 178 |
| google/gemini-3.8-flash | 3 of 3 | 0 / 564 | 1 of 3 | 1 / 98 | 0.59, 0.40 | 1319 / 3853 |
| anthropic/claude-haiku-4.5 | 3 of 3 | 0 / 609 | 3 of 3 | 0 / 12 | 0.89, 0.96 | 1026 / 1574 |
| openai/gpt-6-luna | 3 of 3 | 0 / 848 | 1 of 3 | 0 / 109 | 0.30, 0.08 | 3195 / 5911 |
Lockstep waits for every decision, so it measures the decisions alone; realtime keeps the car moving while a request is in flight, as the demo does, so a slow answer arrives too late. Gemini and Luna lose drives in realtime to latency, not to judgement: about a second per decision, and the demo's 1.5 s timeout hands most of Luna's decisions to the rules driver. Over all of their drives, in lockstep and in realtime:
- the Qwen base passed 16 of 18 and 9 of 9 (in lockstep one collision in s2-4 and one drive off the road in s2-5; the rules driver alone passes both);
- Gemma 4 12B passed 18 of 18 and 9 of 9;
- Gemma 4 31B passed 18 of 18 and 8 of 9 (in realtime a collision in s2-2);
- Cygnet passed 6 of 6 and 9 of 9.
Pong (5 games of up to 45 s, seeds 20260917 to 20260921)
| player | agrees with the oracle while the ball approaches | decisions on approach | decisions per second | latency p50 / p95 | returns per game |
|---|---|---|---|---|---|
| Decisio, Qwen3.6-35B-A3B base | 60% [53, 65] | 531 | 22.9 | 44 / 47 | 22.8 [12.4, 32.8] |
| Decisio, Gemma 4 12B base | 62% [56, 66] | 539 | 22.8 | 48 / 50 | 23.2 [21.2, 24.8] |
| Decisio, Gemma 4 31B base | 85% [82, 88] | 1679 | 15.0 | 68 / 69 | 84.2 [83.0, 86.6] |
| Cygnet | 47% [41, 53] | 459 | 22.9 | 41 / 69 | 19.2 [12.8, 26.4] |
| Jev 1.13 through its public API | 96% [94, 98] | 913 | 8.2 | 116 / 166 | 45.8 [45.4, 46.0] |
| google/gemini-3.8-flash | 98% [94, 100] | 95 | 0.87 | 1131 / 1570 | 5.0 |
| anthropic/claude-haiku-4.5 | 69% [55, 86] | 139 | 1.29 | 745 / 1000 | 6.8 [6.4, 7.0] |
| openai/gpt-6-luna | 98% [95, 100] | 58 | 0.62 | 1460 / 2724 | 3.4 [3.0, 3.8] |
A game ends when the model's paddle misses five times; the ball advances one step per decision, so returns per game reward fast answers as much as right ones. Gemma 4 31B played all five games to the 45 s limit, which is why it was asked three times as often as the other bases.
Triage (400 held-out tickets over 20 queues)
| base | plain question | after registering 200 labelled tickets | paired difference | macro-F1, plain / registered | latency p50, plain / registered |
|---|---|---|---|---|---|
| Qwen3.6-35B-A3B | 91.0% [88.2, 93.8] | 92.5% [89.8, 95.0] | +1.5 points [+0.2, +3.0]; 7 answers fixed, 1 broken | 0.912 / 0.926 | 42 / 120 |
| Gemma 4 12B | 93.8% [91.2, 96.0] | 95.2% [93.0, 97.2] | +1.5 points [+0.0, +3.0]; 8 fixed, 2 broken | 0.938 / 0.952 | 50 / 145 |
| Gemma 4 31B | 97.0% [95.2, 98.5] | 97.2% [95.5, 98.8] | +0.2 points [-0.5, +1.2]; 2 fixed, 1 broken | 0.970 / 0.973 | 71 / 262 |
Registration was repeated on each base and took 33 s, 35 s and 69 s. On the Qwen base it kept both the intent head and the calibration; on the Gemma bases it kept the head only (each is kept only when cross-validation on the examples shows a gain). The head costs latency: a registered answer reads the model's hidden state as well as its letters. The plain question is already right on 91% to 97% of these synthetic tickets, so the lift is small, and smallest where the plain answer is best.
Source: docs/demos/README.md
How each decision is judged
The oracles, the error analysis and the renderings tested on held-out data: the demos, judged decision by decision.
Each base ran from tag v0.6.0 (0ae97ca) as --base <key> --pad-policy row: qwen3.6-35b-a3b, gemma-4-12b and gemma-4-31b, in sequence in one session on one NVIDIA RTX PRO 6000 Blackwell Workstation Edition (600 W) on 2026-10-05, with the demo client on the same machine and the same seeds, oracles and renderings for each.
--pad-policy row pads the state to the base's block: the Qwen base has one (1,056 tokens), and the Gemma bases have none, so on them the flag has no effect.
It is stated under every clip and table.
Cygnet and the hosted models were measured on 2026-10-03, Cygnet on a card of the same type (500 W).
The Qwen base was also the first round's player on 2026-10-03, served from main at 62dfe1f with the same prompt layout. Its counts in Pong, lockstep driving, the browser agent and triage are identical to that round's, and its latencies moved by a few ms; realtime driving differs, because there the latency decides which moments are asked about.