Skip to content

Demos, judged decision by decision

Decisio served the demos in examples/demos/ (Pong, a driving simulator, a browser agent, and a support-ticket triage feed) on each of its three bases, and every decision it made was checked against an oracle. Cygnet and four hosted models played the same games on the same seeds for comparison. This page has the results, the error analysis behind them, the renderings that were tested on held-out data, and what teaching (task registration) did.

Each base ran from tag v0.6.0 (0ae97ca) as --base <key> --pad-policy row: qwen3.6-35b-a3b, gemma-4-12b and gemma-4-31b, in sequence in one session on one NVIDIA RTX PRO 6000 Blackwell Workstation Edition (600 W) on 2026-10-05, with the demo client on the same machine and the same seeds, oracles and renderings for each. --pad-policy row pads the state to the base's block: the Qwen base has one (1,056 tokens), and the Gemma bases have none, so on them the flag has no effect. It is stated under every clip and table. Cygnet and the hosted models were measured on 2026-10-03, Cygnet on a card of the same type (500 W).

The Qwen base was also the first round's player on 2026-10-03, served from main at 62dfe1f with the same prompt layout. Its counts in Pong, lockstep driving, the browser agent and triage are identical to that round's, and its latencies moved by a few ms; realtime driving differs, because there the latency decides which moments are asked about.

Every clip is rendered from a recorded trajectory (examples/demos/tools/trajectory.py), at the speed the run happened. Nothing in a clip is re-simulated or re-driven; the caption under it names the model, the card, the median latency of that run and the padding. Every clip ends on a scoreboard card with that demo's totals for every player, the clip's own player highlighted.

demoQwen3.6-35B-A3B baseGemma 4 12B baseGemma 4 31B base
Pong, the three bases on the same serve (seed 20260918)MP4 · GIFthe same clipthe same clip
Driving, one drive in real time (scenario s1-1)MP4MP4MP4 · GIF (an excerpt)
Browser agent, the travel taskMP4 · GIFMP4 · GIFMP4 · GIF
Browser agent, the reading roomMP4MP4MP4
Triage, the same tickets before and after registering 200 labelled ticketsMP4MP4 · GIF (an excerpt)MP4
demoCygnet beside the Qwen base
Driving, the same drive side by sideMP4
Browser agent, the travel task side by sideMP4

The hosted players' clips are rendered from the same final runs as their table rows, on the same seeds as Decisio's clips, at the speed the run happened, so their latency shows. They show each player's moves only: every option is drawn as chosen or not, with no probability. Jev 1.13 returns a probability for every option; its clips show its moves and not its probabilities. Each caption carries the exact model ID, the date, the route and serving provider, the run's median latency and how many requests had no answer in time, and says "model output" and "AI-generated", rendered by the decisio project.

demoJev 1.13 through its public APIanthropic/claude-haiku-4.5openai/gpt-6-luna
Pong (seed 20260918)MP4MP4MP4
Driving, one drive in real time (scenario s1-1)MP4MP4MP4
Browser agent, the travel taskMP4MP4MP4

There is no clip of google/gemini-3.8-flash: Google's Gemini API terms forbid publicly displaying the content it returns, so Gemini appears in the tables and on the scoreboard cards only. The browser clips hold each step for at least 0.8 s so it can be read; their captions say so. The driving clips draw map data from OpenStreetMap (© OpenStreetMap contributors, ODbL).

demothe oraclewhat it is not
Ponga perfect-information paddle: the true ball's crossing of the paddle's plane after wall bounces, and a move toward it, staying inside half a step (6 units)while the ball moves away there is nothing to judge; those decisions are counted apart
drivingthe demo's own rules driver on the same snapshot and the same candidates: the motion (drive or stop) and the manoeuvrea good driver, not a perfect one; the drive's own score (arrive with no collision, red light, rolled stop sign, failure to yield or a second off the road) is reported beside it
browser agentthe set of actions a careful user could take at each step of two fixture tasks; 272 labelled steps from three placesa label for the operation and the target on two small fixtures
triagethe ticket's true queuesynthetic, templated tickets

A disagreement is classed, in this order:

  1. Latency effect: the right decision arrived too late (driving only: the answer equals the oracle on the snapshot it was asked about but not when it was applied, or the demo's 1.5 s timeout replaced it).
  2. Prompt or state effect: a written detector per demo, defined before any fix was tested: the model followed the question's own wording where it differs from the oracle (Pong's 5-unit band against the game's 6; the driving state's "at" a light up to 6 m before the line, where the rules driver stops within 1.5 m, counted "as worded" and not as an error), or the state lacked what the decision needs (the browser agent: DONE while a typed search is still unsubmitted).
  3. Near-tie: the chosen option's probability minus the oracle option's is below 0.10.
  4. Judgement error: anything else; reported, not tuned away.

The hosted models answer with one option, not a distribution, so their misses cannot be split into near-tie and judgement.

Decisio's three bases and Cygnet ran on the card; the hosted models ran from a Mac on the same seeds and renderings. Driving compares every player on the same three drives (suite seed 1, scenarios 1 to 3); the card players also drove more, given below each table. Intervals are 95% bootstrap intervals over games or drives; latency is the client's round trip, p50 / p95, in ms.

rowmodel IDdaterouteserving provider
Decisio, Qwen3.6-35B-A3B baseQwen/Qwen3.6-35B-A3B-FP8 @95a723d0 served by decisio v0.6.0 (--base qwen3.6-35b-a3b --pad-policy row)2026-10-05local server on the card
Decisio, Gemma 4 12B basegoogle/gemma-4-12B-it @707f0a3b served by decisio v0.6.0 (--base gemma-4-12b --pad-policy row)2026-10-05local server on the card
Decisio, Gemma 4 31B basegoogle/gemma-4-31B-it @842da379, FP8 on load, served by decisio v0.6.0 (--base gemma-4-31b --pad-policy row)2026-10-05local server on the card
Cygnetgoogle/gemma-4-12B-it @707f0a3b with the cygnet-recipe decision server @3cf591c6 (T 3.4)2026-10-03local server on the card
Jev 1.13 through its public APIjev-1.13.02026-10-03TypeSafe public APITypeSafe
google/gemini-3.8-flashgoogle/gemini-3.8-flash2026-10-03OpenRouterGoogle AI Studio
anthropic/claude-haiku-4.5anthropic/claude-haiku-4.52026-10-03OpenRouterAnthropic
openai/gpt-6-lunaopenai/gpt-6-luna2026-10-03OpenRouterOpenAI

The hosted models' latencies include the network from the Mac; OpenRouter's own added time was measured at 130 to 155 ms per call.

playertravel completedtravel decisions rightreading room completedreading room decisions rightlatency p50 / p95
Decisio, Qwen3.6-35B-A3B base10/1050/7010/1080/140141 / 174
Decisio, Gemma 4 12B base10/1060/6010/1020/20141 / 275
Decisio, Gemma 4 31B base10/1040/5010/1020/20184 / 417
Cygnet10/1060/6010/1020/2058 / 253
Jev 1.13 through its public API10/1060/6010/1020/20130 / 185
google/gemini-3.8-flash10/1060/6010/1020/201308 / 2030
anthropic/claude-haiku-4.510/1040/5010/1020/201248 / 1767
openai/gpt-6-luna10/1060/6810/1020/202878 / 4464

On the 272 labelled steps (the operation question alone), the answer is inside the acceptable set on 243 (89%) for the Qwen base, 269 (99%) for Gemma 4 12B, 248 (91%) for Gemma 4 31B and 268 (99%) for Cygnet.

playerlockstep: drives passedlockstep: motion errors / motion answersrealtime: drives passedrealtime: motion errors / latency effectsdecisions per second (lockstep, realtime)latency p50 / p95 (lockstep)
Decisio, Qwen3.6-35B-A3B base3 of 38 / 5453 of 35 / 27.74, 3.12102 / 158
Decisio, Gemma 4 12B base3 of 325 / 7923 of 36 / 44.01, 2.99218 / 343
Decisio, Gemma 4 31B base3 of 35 / 5683 of 30 / 32.70, 2.37341 / 569
Cygnet3 of 30 / 5743 of 30 / 72.49, 2.30371 / 693
Jev 1.13 through its public API3 of 30 / 7113 of 30 / 26.58, 3.03126 / 178
google/gemini-3.8-flash3 of 30 / 5641 of 31 / 980.59, 0.401319 / 3853
anthropic/claude-haiku-4.53 of 30 / 6093 of 30 / 120.89, 0.961026 / 1574
openai/gpt-6-luna3 of 30 / 8481 of 30 / 1090.30, 0.083195 / 5911

Lockstep waits for every decision, so it measures the decisions alone; realtime keeps the car moving while a request is in flight, as the demo does, so a slow answer arrives too late. Gemini and Luna lose drives in realtime to latency, not to judgement: about a second per decision, and the demo's 1.5 s timeout hands most of Luna's decisions to the rules driver. Over all of their drives, in lockstep and in realtime:

  • the Qwen base passed 16 of 18 and 9 of 9 (in lockstep one collision in s2-4 and one drive off the road in s2-5; the rules driver alone passes both);
  • Gemma 4 12B passed 18 of 18 and 9 of 9;
  • Gemma 4 31B passed 18 of 18 and 8 of 9 (in realtime a collision in s2-2);
  • Cygnet passed 6 of 6 and 9 of 9.

Pong (5 games of up to 45 s, seeds 20260917 to 20260921)

Section titled “Pong (5 games of up to 45 s, seeds 20260917 to 20260921)”
playeragrees with the oracle while the ball approachesdecisions on approachdecisions per secondlatency p50 / p95returns per game
Decisio, Qwen3.6-35B-A3B base60% [53, 65]53122.944 / 4722.8 [12.4, 32.8]
Decisio, Gemma 4 12B base62% [56, 66]53922.848 / 5023.2 [21.2, 24.8]
Decisio, Gemma 4 31B base85% [82, 88]167915.068 / 6984.2 [83.0, 86.6]
Cygnet47% [41, 53]45922.941 / 6919.2 [12.8, 26.4]
Jev 1.13 through its public API96% [94, 98]9138.2116 / 16645.8 [45.4, 46.0]
google/gemini-3.8-flash98% [94, 100]950.871131 / 15705.0
anthropic/claude-haiku-4.569% [55, 86]1391.29745 / 10006.8 [6.4, 7.0]
openai/gpt-6-luna98% [95, 100]580.621460 / 27243.4 [3.0, 3.8]

A game ends when the model's paddle misses five times; the ball advances one step per decision, so returns per game reward fast answers as much as right ones. Gemma 4 31B played all five games to the 45 s limit, which is why it was asked three times as often as the other bases.

Triage (400 held-out tickets over 20 queues)

Section titled “Triage (400 held-out tickets over 20 queues)”
baseplain questionafter registering 200 labelled ticketspaired differencemacro-F1, plain / registeredlatency p50, plain / registered
Qwen3.6-35B-A3B91.0% [88.2, 93.8]92.5% [89.8, 95.0]+1.5 points [+0.2, +3.0]; 7 answers fixed, 1 broken0.912 / 0.92642 / 120
Gemma 4 12B93.8% [91.2, 96.0]95.2% [93.0, 97.2]+1.5 points [+0.0, +3.0]; 8 fixed, 2 broken0.938 / 0.95250 / 145
Gemma 4 31B97.0% [95.2, 98.5]97.2% [95.5, 98.8]+0.2 points [-0.5, +1.2]; 2 fixed, 1 broken0.970 / 0.97371 / 262

Registration was repeated on each base and took 33 s, 35 s and 69 s. On the Qwen base it kept both the intent head and the calibration; on the Gemma bases it kept the head only (each is kept only when cross-validation on the examples shows a gain). The head costs latency: a registered answer reads the model's hidden state as well as its letters. The plain question is already right on 91% to 97% of these synthetic tickets, so the lift is small, and smallest where the plain answer is best.

demobasedecisions judgedagreelatencyprompt or statenear-tiejudgementwhat the errors are
Pong, ball approachingQwen3.6-35B-A3B5313300 by construction1263126mostly "stay" where the perfect paddle moves down (85 of 201); 21 of the misses turned a return into a lost point
Gemma 4 12B5393320 by construction413190mostly "down" where the perfect paddle moves up (97 of 207); 24 turned a return into a lost point
Gemma 4 31B1,6791,4210 by construction136239mostly "stay" where the perfect paddle moves down (210 of 258); across its misses the ball crosses a median 7 units from the paddle, just outside the 6-unit stay band; none cost a point
driving lockstep, motion (18 drives)Qwen3.6-35B-A3B4,1732,934 + 1,207 as worded0131930 of 32 a stop where the rules driver goes
Gemma 4 12B5,0613,927 + 1,108 as worded0125all 26 a stop where the rules driver goes
Gemma 4 31B4,2533,055 + 1,182 as worded001614 of 16 a stop where the rules driver goes
driving realtime, motion (9 drives)Qwen3.6-35B-A3B1,6971,077 + 607 as worded715all 6 a stop where the rules driver goes
Gemma 4 12B2,0871,471 + 598 as worded1206all 6 a stop where the rules driver goes
Gemma 4 31B1,554836 + 702 as worded1213all 4 a stop where the rules driver goes
browser, labelled stepsQwen3.6-35B-A3B2722430 by construction5420DONE before the search is applied; a click on the wrong control
Gemma 4 12B2722690 by construction012
Gemma 4 31B2722480 by construction1707mostly DONE while a typed search is still unsubmitted

In the live browser runs every base completes both tasks every time, but not always by the shortest path. On the travel task the Qwen base re-selects the category once and opens Casa Flora before submitting the typed destination, and Gemma 4 31B opens Casa Flora without submitting it; the fixture's own check does not test the search, so the runs count as completed and the oracle counts those steps as misses. In the reading room the Qwen base opens the article and then clicks back to the list about six times (its probability for DONE on the article is 0.27 to 0.42) before DONE wins. Gemma 4 12B takes the shortest path in both tasks in every run, and Gemma 4 31B in the reading room.

These tests ran on the Qwen base in the first round, before the bases session; they were not repeated on the Gemma bases. A change to a demo's state or question is kept only when its gain on held-out seeds or places is clear of zero (the paired 95% interval excludes it) and it adds observations only: no verdict, no recommendation, no score per option.

demochangeheld-out dataresultkept
Pongthe arrival in words ("the ball will arrive 14.2 units above the centre of your paddle") instead of coordinates, and the moves defined ("move the paddle 12 units toward the top of the screen")seeds 40000 to 40009-32.3 points [-40.8, -23.9]: Decisio answers "stay" almost alwaysno
Pongoffset, the intercept minus the paddle's position, added to the statethe same+4.9 points [-4.9, +15.3]no
Pongthe stay band worded as the game's 6 unitsthe same-1.9 points [-11.4, +9.8]no
drivingafter a completed stop, once the car is past the line, the state says "the car is already in the junction, N m past the stop line" instead of telling it to hold for cross trafficsuite seeds 201 to 203, 18 drives, lockstepmotion errors 101 to 59, -2.33 per drive [-5.17, -0.44]; drives passed 14 of 18 bothyes
browser agenta typed field marked submitted or not (from the page's own submit events), filters marked applied, and a shorter action historythe 272 labelled steps of three places, and 10 + 10 live runsall steps -3.3 points [-7.4, +0.7] (Casa Flora +0.8, the Glasshouse -8.3 [-15.0, -1.7]); the reading room then loops to the step limit in 10 of 10 runsno
browser agentDONE's description adds that an unsubmitted search is not appliedthe labelled steps of the two places it was not designed on (it came from a Casa Flora step)the Glasshouse and Serra Lodge 88% to 86% and 91% to 88%; Casa Flora itself +5.0 points [+0.8, +10.0]; all 272 steps +0.7 [-2.9, +4.0]no

The driving change fixes a sentence that contradicted the state it sits in: the default text tells the car to hold for cross traffic even when the car is already in the junction, where the rules driver goes on. The final runs above use it; Pong and the browser agent use their default renderings.

With the default prompt the Pong question asks the model to compare two numbers, the intercept and the paddle's position, and choose a direction. The Qwen base agrees with the perfect paddle on 60% of decisions while the ball approaches, on the evaluated seeds and on held-out seeds alike; given the signed distance it reaches 66% and reads the sign backwards on most of what remains, and given the relationship in words it stops moving. Gemma 4 12B is at 62%, and Gemma 4 31B reaches 85%, its misses mostly a "stay" close to the edge of the stay band. This is recorded as a limit of the models on numeric comparison, not of the prompt, and the larger model lifts it. The same question is answered at 96% by Jev 1.13 through its public API in about 120 ms per decision, and at 98% by Gemini 3.8 Flash and GPT 6 Luna in 1.1 to 1.5 s.

In the games, on the Qwen base in the first round, registering the questions with fixed option lists (Pong's move, the driving motion, the browser operation) from oracle-labelled examples on seeds disjoint from the evaluated ones gave no gain. The server declined the Pong and driving calibrations because cross-validation on their examples showed none, so those questions were answered exactly as before. It kept the browser operation's calibration, which moved probability toward the options the examples used most and scored -1.8 points [-4.4, +0.4] on the labelled steps, with the live runs unchanged; seven per-option biases cannot make a correction that depends on the page. The driving manoeuvre and the browser's element target cannot be registered: their options change every step. On triage, a fixed 20-queue question, registration from 10 labelled tickets per queue lifted held-out accuracy by 1.5 points on the Qwen base ([+0.2, +3.0], from 91.0%) and on Gemma 4 12B ([+0.0, +3.0], from 93.8%), and by 0.2 points [-0.5, +1.2] on Gemma 4 31B, from 97.0%.

examples/demos/README.md lists the commands: each tools/measure_*.py writes a run record with trajectories, and each tools/render_*.py draws a clip from them. The hosted models were reached through examples/demos/tools/systemone_gateway.py, which renders a System One request as a chat request with a structured answer; the OpenRouter calls pinned the first-party provider, disallowed fallbacks and denied data collection.

Decisio is an independent open-source project, not affiliated with or endorsed by TypeSafe. Jev is TypeSafe's product name.

Apache-2.0. This page is built from Decisio v0.8.2.