Demos

Decisio served the demos in examples/demos/ (Pong, a driving simulator, a browser agent, and a support-ticket triage feed) on each of its three bases, and every decision it made was checked against an oracle. Cygnet and four hosted models played the same games on the same seeds for comparison. This page has the results, the error analysis behind them, the renderings that were tested on held-out data, and what teaching (task registration) did.

Watch it play

Every clip is rendered from a recorded trajectory (examples/demos/tools/trajectory.py), at the speed the run happened. Nothing in a clip is re-simulated or re-driven; the caption under it names the model, the card, the median latency of that run and the padding. Every clip ends on a scoreboard card with that demo's totals for every player, the clip's own player highlighted.

Pong, the three bases on the same serve (seed 20260918)

Pong, the three bases on the same serve (seed 20260918)

Driving, one drive in real time (scenario s1-1)

Qwen3.6-35B-A3B base
Gemma 4 12B base
Gemma 4 31B base

Browser agent, the travel task

Qwen3.6-35B-A3B base
Gemma 4 12B base
Gemma 4 31B base

Browser agent, the reading room

Qwen3.6-35B-A3B base
Gemma 4 12B base
Gemma 4 31B base

Triage, the same tickets before and after registering 200 labelled tickets

Qwen3.6-35B-A3B base
Gemma 4 12B base
Gemma 4 31B base

Driving, the same drive side by side

Driving, the same drive side by side

Browser agent, the travel task side by side

Browser agent, the travel task side by side

Jev 1.13 through its public API, anthropic/claude-haiku-4.5 and openai/gpt-6-luna played the same seeds; their clips are on the repository's demos page.

Results

Decisio's three bases and Cygnet ran on the card; the hosted models ran from a Mac on the same seeds and renderings. Driving compares every player on the same three drives (suite seed 1, scenarios 1 to 3); the card players also drove more, given below each table. Intervals are 95% bootstrap intervals over games or drives; latency is the client's round trip, p50 / p95, in ms.

The players

The players
rowmodel IDdaterouteserving provider
Decisio, Qwen3.6-35B-A3B baseQwen/Qwen3.6-35B-A3B-FP8 @95a723d0 served by decisio v0.6.0 (--base qwen3.6-35b-a3b --pad-policy row)2026-10-05local server on the card
Decisio, Gemma 4 12B basegoogle/gemma-4-12B-it @707f0a3b served by decisio v0.6.0 (--base gemma-4-12b --pad-policy row)2026-10-05local server on the card
Decisio, Gemma 4 31B basegoogle/gemma-4-31B-it @842da379, FP8 on load, served by decisio v0.6.0 (--base gemma-4-31b --pad-policy row)2026-10-05local server on the card
Cygnetgoogle/gemma-4-12B-it @707f0a3b with the cygnet-recipe decision server @3cf591c6 (T 3.4)2026-10-03local server on the card
Jev 1.13 through its public APIjev-1.13.02026-10-03TypeSafe public APITypeSafe
google/gemini-3.8-flashgoogle/gemini-3.8-flash2026-10-03OpenRouterGoogle AI Studio
anthropic/claude-haiku-4.5anthropic/claude-haiku-4.52026-10-03OpenRouterAnthropic
openai/gpt-6-lunaopenai/gpt-6-luna2026-10-03OpenRouterOpenAI

The hosted models' latencies include the network from the Mac; OpenRouter's own added time was measured at 130 to 155 ms per call.

Browser agent (10 runs of each task)

Browser agent (10 runs of each task)
playertravel completedtravel decisions rightreading room completedreading room decisions rightlatency p50 / p95
Decisio, Qwen3.6-35B-A3B base10/1050/7010/1080/140141 / 174
Decisio, Gemma 4 12B base10/1060/6010/1020/20141 / 275
Decisio, Gemma 4 31B base10/1040/5010/1020/20184 / 417
Cygnet10/1060/6010/1020/2058 / 253
Jev 1.13 through its public API10/1060/6010/1020/20130 / 185
google/gemini-3.8-flash10/1060/6010/1020/201308 / 2030
anthropic/claude-haiku-4.510/1040/5010/1020/201248 / 1767
openai/gpt-6-luna10/1060/6810/1020/202878 / 4464

On the 272 labelled steps (the operation question alone), the answer is inside the acceptable set on 243 (89%) for the Qwen base, 269 (99%) for Gemma 4 12B, 248 (91%) for Gemma 4 31B and 268 (99%) for Cygnet.

Driving

Driving
playerlockstep: drives passedlockstep: motion errors / motion answersrealtime: drives passedrealtime: motion errors / latency effectsdecisions per second (lockstep, realtime)latency p50 / p95 (lockstep)
Decisio, Qwen3.6-35B-A3B base3 of 38 / 5453 of 35 / 27.74, 3.12102 / 158
Decisio, Gemma 4 12B base3 of 325 / 7923 of 36 / 44.01, 2.99218 / 343
Decisio, Gemma 4 31B base3 of 35 / 5683 of 30 / 32.70, 2.37341 / 569
Cygnet3 of 30 / 5743 of 30 / 72.49, 2.30371 / 693
Jev 1.13 through its public API3 of 30 / 7113 of 30 / 26.58, 3.03126 / 178
google/gemini-3.8-flash3 of 30 / 5641 of 31 / 980.59, 0.401319 / 3853
anthropic/claude-haiku-4.53 of 30 / 6093 of 30 / 120.89, 0.961026 / 1574
openai/gpt-6-luna3 of 30 / 8481 of 30 / 1090.30, 0.083195 / 5911

Lockstep waits for every decision, so it measures the decisions alone; realtime keeps the car moving while a request is in flight, as the demo does, so a slow answer arrives too late. Gemini and Luna lose drives in realtime to latency, not to judgement: about a second per decision, and the demo's 1.5 s timeout hands most of Luna's decisions to the rules driver. Over all of their drives, in lockstep and in realtime:

  • the Qwen base passed 16 of 18 and 9 of 9 (in lockstep one collision in s2-4 and one drive off the road in s2-5; the rules driver alone passes both);
  • Gemma 4 12B passed 18 of 18 and 9 of 9;
  • Gemma 4 31B passed 18 of 18 and 8 of 9 (in realtime a collision in s2-2);
  • Cygnet passed 6 of 6 and 9 of 9.

Pong (5 games of up to 45 s, seeds 20260917 to 20260921)

Pong (5 games of up to 45 s, seeds 20260917 to 20260921)
playeragrees with the oracle while the ball approachesdecisions on approachdecisions per secondlatency p50 / p95returns per game
Decisio, Qwen3.6-35B-A3B base60% [53, 65]53122.944 / 4722.8 [12.4, 32.8]
Decisio, Gemma 4 12B base62% [56, 66]53922.848 / 5023.2 [21.2, 24.8]
Decisio, Gemma 4 31B base85% [82, 88]167915.068 / 6984.2 [83.0, 86.6]
Cygnet47% [41, 53]45922.941 / 6919.2 [12.8, 26.4]
Jev 1.13 through its public API96% [94, 98]9138.2116 / 16645.8 [45.4, 46.0]
google/gemini-3.8-flash98% [94, 100]950.871131 / 15705.0
anthropic/claude-haiku-4.569% [55, 86]1391.29745 / 10006.8 [6.4, 7.0]
openai/gpt-6-luna98% [95, 100]580.621460 / 27243.4 [3.0, 3.8]

A game ends when the model's paddle misses five times; the ball advances one step per decision, so returns per game reward fast answers as much as right ones. Gemma 4 31B played all five games to the 45 s limit, which is why it was asked three times as often as the other bases.

Triage (400 held-out tickets over 20 queues)

Triage (400 held-out tickets over 20 queues)
baseplain questionafter registering 200 labelled ticketspaired differencemacro-F1, plain / registeredlatency p50, plain / registered
Qwen3.6-35B-A3B91.0% [88.2, 93.8]92.5% [89.8, 95.0]+1.5 points [+0.2, +3.0]; 7 answers fixed, 1 broken0.912 / 0.92642 / 120
Gemma 4 12B93.8% [91.2, 96.0]95.2% [93.0, 97.2]+1.5 points [+0.0, +3.0]; 8 fixed, 2 broken0.938 / 0.95250 / 145
Gemma 4 31B97.0% [95.2, 98.5]97.2% [95.5, 98.8]+0.2 points [-0.5, +1.2]; 2 fixed, 1 broken0.970 / 0.97371 / 262

Registration was repeated on each base and took 33 s, 35 s and 69 s. On the Qwen base it kept both the intent head and the calibration; on the Gemma bases it kept the head only (each is kept only when cross-validation on the examples shows a gain). The head costs latency: a registered answer reads the model's hidden state as well as its letters. The plain question is already right on 91% to 97% of these synthetic tickets, so the lift is small, and smallest where the plain answer is best.

Source: docs/demos/README.md

How each decision is judged

The oracles, the error analysis and the renderings tested on held-out data: the demos, judged decision by decision.

Each base ran from tag v0.6.0 (0ae97ca) as --base <key> --pad-policy row: qwen3.6-35b-a3b, gemma-4-12b and gemma-4-31b, in sequence in one session on one NVIDIA RTX PRO 6000 Blackwell Workstation Edition (600 W) on 2026-10-05, with the demo client on the same machine and the same seeds, oracles and renderings for each. --pad-policy row pads the state to the base's block: the Qwen base has one (1,056 tokens), and the Gemma bases have none, so on them the flag has no effect. It is stated under every clip and table. Cygnet and the hosted models were measured on 2026-10-03, Cygnet on a card of the same type (500 W).

The Qwen base was also the first round's player on 2026-10-03, served from main at 62dfe1f with the same prompt layout. Its counts in Pong, lockstep driving, the browser agent and triage are identical to that round's, and its latencies moved by a few ms; realtime driving differs, because there the latency decides which moments are asked about.