Choosing a base
Three base models are served behind the same routes, wire format and features, one per server, chosen with --base:
uv run python -m decisio.serve.vllm_engine --base qwen3.6-35b-a3b # the defaultuv run python -m decisio.serve.vllm_engine --base gemma-4-12buv run python -m decisio.serve.vllm_engine --base gemma-4-31bNeeds an NVIDIA card: taken verbatim from the repository's tested docs, run record runs/2026-10-04_gemma-base/.
--base brings its checkpoint at a pinned revision and every setting below; a flag given explicitly overrides the base's value.
Each base's settings were measured as one configuration.
| Base | Checkpoint | Provenance |
|---|---|---|
qwen3.6-35b-a3b (default) | Qwen/Qwen3.6-35B-A3B-FP8 at 95a723d0, 33.3 GiB in memory | Official checkpoint from Alibaba's Qwen team, at a pinned revision; no adapter or fine-tuning by us |
gemma-4-12b | google/gemma-4-12B-it at 707f0a3b, bf16, 22.8 GiB in memory | Official checkpoint from Google, at a pinned revision; no adapter or fine-tuning by us |
gemma-4-31b | google/gemma-4-31B-it at 842da379, quantized to FP8 when it loads (vLLM 0.30.0), 30.6 GiB in memory | Official checkpoint from Google, at a pinned revision; no adapter or fine-tuning by us; FP8 on load, which a pinned FP8 checkpoint replaces when one exists |
A base that is a third-party fine-tune names its publisher in the provenance column and states what it was trained on.
| Setting | qwen3.6-35b-a3b (default) | gemma-4-12b | gemma-4-31b |
|---|---|---|---|
| Temperatures | 1.370 for choice questions, 1.506 for yes/no and score | 3.592 for every question type | 4.672 for choice questions, 5.252 for yes/no and score |
| Prompt | no system turn, read after an "Answer:" prefill, one token per option letter | a system turn, read at the chat template's own answer position, every single-token form of each letter summed | as gemma-4-12b |
| Yes/no | a two-option letter choice with its sides named | a two-option letter choice, each side shown as its description | as gemma-4-12b |
| Padding | the state padded to the 1,056-token block | none | none |
| Several questions in one request | the same probabilities as each question sent alone, bit for bit | the same choice as each question sent alone, probabilities within 0.035 | scored in one batch after the state is read (--multi-question warm, its default since 0.8.0): the same choice as one engine call per question on every item measured, probabilities within 0.0073 |
| Precision | FP8, as the checkpoint stores it | bf16 | FP8 on load (at bf16 it leaves too little of a 96 GB card for a 32,768-token context) |
| A second, different question on a document already read (since 0.8.1; one RTX PRO 6000 at 585 W, AMD Ryzen Threadripper 9960X) | read from the cache, as always: 23.5 ms after an 85 ms first read (3,000 tokens) | read from the cache since 0.8.1: 35.8 ms after a 260 ms first read; the first read of a new document costs +16 to +28 ms more (300 to 3,000 tokens) | read from the cache since 0.8.1: 43.3 ms after a 439 ms first read; the first read costs +21 to +34 ms more |
In plain words:
- Choose Gemma 4 12B for committed yes/no answers, scores and intent routing on taxonomies like CLINC150.
- Choose Qwen3.6-35B-A3B for knowledge questions, long states seen for the first time, and answers that repeat bit for bit.
- On JevBench's published items they are level.
- Gemma 4 12B does not fit a 32 GB card on vLLM; on a Mac it runs from its MLX conversion (Mac with MLX).
- Choose Gemma 4 31B for knowledge questions and the strongest suite and JevBench readings: it is stronger than both served bases on every accuracy measure we have, and slower on states it has not seen before, the more so the longer the state; its weights take 31 GB, and it needs a 96 GB card on vLLM (no Mac build).
What drives the choice, measured on one RTX PRO 6000 Blackwell in one session, each base with its own defaults, paired over the same items (95% bootstrap intervals; runs/2026-10-04_gemma-base/):
| Measure | Gemma 4 12B | Qwen3.6-35B-A3B | Gemma minus Qwen |
|---|---|---|---|
| JevBench v1.5 open-set reading, I_open (equal types) | 64.4 | 49.4 | +15.0 [+7.0, +23.6] |
| JevBench v1.5, yes/no / score | 49.0 / 63.7 | 13.7 / 57.2 | |
| Yes/no answers between 0.20 and 0.80 (74 items) | 15% | 39% | |
| JevBench, published items, correct | 0.0 points [-4.3, +4.3] | ||
| Decision Index accuracy, CLINC150+OOS | +4.5 [+3.6, +5.5] | ||
| Decision Index accuracy, BANKING77 | -1.4 [-2.7, -0.1] | ||
| Decision Index accuracy, GPQA Diamond | -13.6 [-21.7, -5.6] | ||
| Decision Index accuracy, MMLU-Pro | -6.4 [-7.2, -5.5] | ||
| 1,400-item suite, accuracy | 0.735 | 0.770 | -3.5 [-5.5, -1.6] |
One question on a new 3,000-token state, server time (0.8.1's served defaults, the engine in the server's process; one RTX PRO 6000 at 585 W, AMD Ryzen Threadripper 9960X; one session, runs/2026-10-06_latency-585w/) | 258.6 ms | 85.2 ms |
The full paired table, the Gemma base's repeatability and its limits are in EVAL_CARD.md section 6.
Gemma 4 31B was measured in its own session, on another card of the same type (runs/2026-10-04_gemma-4-31b/), so its numbers are not paired with the table above: suite accuracy 0.799 (Qwen 0.770, Gemma 4 12B 0.735), 213 of JevBench's 231 published questions correct (200 each), Decision Index MMLU-Pro 0.694 and GPQA Diamond 0.520 (Qwen 0.613 and 0.510), one question on a new 3,000-token state 438 ms (Qwen 85 ms; 0.8.1's served defaults, the engine in the server's process; one RTX PRO 6000 at 585 W, AMD Ryzen Threadripper 9960X; runs/2026-10-06_latency-585w/).
It went in under the maintainer's decision, past a pre-registered rule it missed by 0.2 to 1.1 points on three of four benchmarks; EVAL_CARD.md section 7 states the rule, the numbers and the reason.
On the four demos, run on each base in one session (docs/demos), Gemma 4 31B played Pong best (85% agreement with a perfect paddle, against 60% and 62%), made the fewest driving motion errors and was most accurate on triage (97.0% against 93.8% for Gemma 4 12B and 91.0% for Qwen), but lost one of nine real-time drives. Gemma 4 12B was the most accurate browser agent (269 of 272 labelled steps, against 248 and 243) and took the shortest path in every run. Qwen3.6-35B-A3B answered fastest in Pong, driving and triage (Pong 44 ms per decision, against 48 and 68 ms; level with Gemma 4 12B in the browser agent), but made the most browser-agent errors.
Decisio is an independent open-source project, not affiliated with or endorsed by TypeSafe. Jev is TypeSafe's product name.
Apache-2.0. This page is built from Decisio v0.8.2.