Skip to content

Choosing a base

Three base models are served behind the same routes, wire format and features, one per server, chosen with --base:

uv run python -m decisio.serve.vllm_engine --base qwen3.6-35b-a3b # the default
uv run python -m decisio.serve.vllm_engine --base gemma-4-12b
uv run python -m decisio.serve.vllm_engine --base gemma-4-31b

Needs an NVIDIA card: taken verbatim from the repository's tested docs, run record runs/2026-10-04_gemma-base/.

--base brings its checkpoint at a pinned revision and every setting below; a flag given explicitly overrides the base's value. Each base's settings were measured as one configuration.

BaseCheckpointProvenance
qwen3.6-35b-a3b (default)Qwen/Qwen3.6-35B-A3B-FP8 at 95a723d0, 33.3 GiB in memoryOfficial checkpoint from Alibaba's Qwen team, at a pinned revision; no adapter or fine-tuning by us
gemma-4-12bgoogle/gemma-4-12B-it at 707f0a3b, bf16, 22.8 GiB in memoryOfficial checkpoint from Google, at a pinned revision; no adapter or fine-tuning by us
gemma-4-31bgoogle/gemma-4-31B-it at 842da379, quantized to FP8 when it loads (vLLM 0.30.0), 30.6 GiB in memoryOfficial checkpoint from Google, at a pinned revision; no adapter or fine-tuning by us; FP8 on load, which a pinned FP8 checkpoint replaces when one exists

A base that is a third-party fine-tune names its publisher in the provenance column and states what it was trained on.

Settingqwen3.6-35b-a3b (default)gemma-4-12bgemma-4-31b
Temperatures1.370 for choice questions, 1.506 for yes/no and score3.592 for every question type4.672 for choice questions, 5.252 for yes/no and score
Promptno system turn, read after an "Answer:" prefill, one token per option lettera system turn, read at the chat template's own answer position, every single-token form of each letter summedas gemma-4-12b
Yes/noa two-option letter choice with its sides nameda two-option letter choice, each side shown as its descriptionas gemma-4-12b
Paddingthe state padded to the 1,056-token blocknonenone
Several questions in one requestthe same probabilities as each question sent alone, bit for bitthe same choice as each question sent alone, probabilities within 0.035scored in one batch after the state is read (--multi-question warm, its default since 0.8.0): the same choice as one engine call per question on every item measured, probabilities within 0.0073
PrecisionFP8, as the checkpoint stores itbf16FP8 on load (at bf16 it leaves too little of a 96 GB card for a 32,768-token context)
A second, different question on a document already read (since 0.8.1; one RTX PRO 6000 at 585 W, AMD Ryzen Threadripper 9960X)read from the cache, as always: 23.5 ms after an 85 ms first read (3,000 tokens)read from the cache since 0.8.1: 35.8 ms after a 260 ms first read; the first read of a new document costs +16 to +28 ms more (300 to 3,000 tokens)read from the cache since 0.8.1: 43.3 ms after a 439 ms first read; the first read costs +21 to +34 ms more

In plain words:

  • Choose Gemma 4 12B for committed yes/no answers, scores and intent routing on taxonomies like CLINC150.
  • Choose Qwen3.6-35B-A3B for knowledge questions, long states seen for the first time, and answers that repeat bit for bit.
  • On JevBench's published items they are level.
  • Gemma 4 12B does not fit a 32 GB card on vLLM; on a Mac it runs from its MLX conversion (Mac with MLX).
  • Choose Gemma 4 31B for knowledge questions and the strongest suite and JevBench readings: it is stronger than both served bases on every accuracy measure we have, and slower on states it has not seen before, the more so the longer the state; its weights take 31 GB, and it needs a 96 GB card on vLLM (no Mac build).

What drives the choice, measured on one RTX PRO 6000 Blackwell in one session, each base with its own defaults, paired over the same items (95% bootstrap intervals; runs/2026-10-04_gemma-base/):

MeasureGemma 4 12BQwen3.6-35B-A3BGemma minus Qwen
JevBench v1.5 open-set reading, I_open (equal types)64.449.4+15.0 [+7.0, +23.6]
JevBench v1.5, yes/no / score49.0 / 63.713.7 / 57.2
Yes/no answers between 0.20 and 0.80 (74 items)15%39%
JevBench, published items, correct0.0 points [-4.3, +4.3]
Decision Index accuracy, CLINC150+OOS+4.5 [+3.6, +5.5]
Decision Index accuracy, BANKING77-1.4 [-2.7, -0.1]
Decision Index accuracy, GPQA Diamond-13.6 [-21.7, -5.6]
Decision Index accuracy, MMLU-Pro-6.4 [-7.2, -5.5]
1,400-item suite, accuracy0.7350.770-3.5 [-5.5, -1.6]
One question on a new 3,000-token state, server time (0.8.1's served defaults, the engine in the server's process; one RTX PRO 6000 at 585 W, AMD Ryzen Threadripper 9960X; one session, runs/2026-10-06_latency-585w/)258.6 ms85.2 ms

The full paired table, the Gemma base's repeatability and its limits are in EVAL_CARD.md section 6.

Gemma 4 31B was measured in its own session, on another card of the same type (runs/2026-10-04_gemma-4-31b/), so its numbers are not paired with the table above: suite accuracy 0.799 (Qwen 0.770, Gemma 4 12B 0.735), 213 of JevBench's 231 published questions correct (200 each), Decision Index MMLU-Pro 0.694 and GPQA Diamond 0.520 (Qwen 0.613 and 0.510), one question on a new 3,000-token state 438 ms (Qwen 85 ms; 0.8.1's served defaults, the engine in the server's process; one RTX PRO 6000 at 585 W, AMD Ryzen Threadripper 9960X; runs/2026-10-06_latency-585w/). It went in under the maintainer's decision, past a pre-registered rule it missed by 0.2 to 1.1 points on three of four benchmarks; EVAL_CARD.md section 7 states the rule, the numbers and the reason.

On the four demos, run on each base in one session (docs/demos), Gemma 4 31B played Pong best (85% agreement with a perfect paddle, against 60% and 62%), made the fewest driving motion errors and was most accurate on triage (97.0% against 93.8% for Gemma 4 12B and 91.0% for Qwen), but lost one of nine real-time drives. Gemma 4 12B was the most accurate browser agent (269 of 272 labelled steps, against 248 and 243) and took the shortest path in every run. Qwen3.6-35B-A3B answered fastest in Pong, driving and triage (Pong 44 ms per decision, against 48 and 68 ms; level with Gemma 4 12B in the browser agent), but made the most browser-agent errors.

Decisio is an independent open-source project, not affiliated with or endorsed by TypeSafe. Jev is TypeSafe's product name.

Apache-2.0. This page is built from Decisio v0.8.2.