Benchmarks

Each base with its own defaults on one RTX PRO 6000 Blackwell: the Qwen and Gemma 4 12B bases in one session (runs/2026-10-04_gemma-base/), Gemma 4 31B in its own (runs/2026-10-04_gemma-4-31b/), image input in another (runs/2026-09-27_image-input/). Measured by us with the public harnesses (the Decision Index kit 0.2.1) and recorded in runs/; none is a board score.

How Decisio compares

decisio's three bases set beside TypeSafe's Jev 1.13 and the three open-weights entries with the highest Decision Index on its public board, as of 2026-10-05. Our values are self-run: the Decision Index runs are submitted in apolinario/decision-index#62 and pending the maintainers' validation, and the JevBench counts are this repository's records, not board entries. Jev's and the other entries' values come only from the public boards, with each board's revision, the fields read and the date read in runs/2026-10-05_comparison/. No figure here comes from any call of ours to TypeSafe's API, and nothing here is a rank.

How a cell is marked. Each of our cells is ahead of Jev, level or behind (bold on this site when level or ahead), by one rule for every cell, written and committed before anything was computed (RULE.md). A difference counts only when it lies outside a 95% interval for the benchmark's size, h = 1.96 √(2m(1 − m)/n), with m the two values' mean and n the benchmark's items; otherwise the cell is level. For the index and the areas, the same per-benchmark variance is carried through the board's formula. The one lower-is-better metric, ForecastBench's Brier loss, is compared in its own direction. Only our cells are marked; the other entries' columns are shown as the board has them.

The headline

Bold: level with or ahead of Jev 1.13, by the comparison's rule; plain: behind.

Decisio's three bases beside Jev 1.13 and the highest open entry, Surogate Rune 26B-A4B v3
MeasureDecisioQwen3.6-35B-A3BDecisioGemma 4 12BDecisioGemma 4 31BJev 1.13Surogate Rune26B-A4B v3
Decision Index 0.2.1 (0 to 100)48.06: behind Jev49.43: behind Jev57.58: level with Jev57.9157.44
Knowledge & Reasoning32.45: behind Jev31.37: behind Jev45.57: behind Jev51.4043.36
Language Understanding51.39: behind Jev51.66: behind Jev61.02: level with Jev62.0263.06
Retrieval & Classification53.56: level with Jev58.00: ahead of Jev62.75: ahead of Jev55.4263.52
Tools & Automation68.43: behind Jev71.62: behind Jev75.04: level with Jev75.0971.22
Arts & Human Taste31.57: behind Jev32.59: behind Jev37.43: level with Jev37.6641.90
On JevBench's 231 published questions, correct200: level with Jev200: level with Jev213: ahead of Jev200not compared (h)
JevBench hard tier, correct82 of 11182 of 11193 of 111not compared (a)not compared (h)
Latency, price or self-hosting cost, context length and what each system is
SystemLatency, median per requestPrice or self-hosting cost, one pass over the 38 index benchmarksContext lengthWhat it is
DecisioQwen3.6-35B-A3B70.5 ms (b)$6.98 (e)32,768 tokens per prompt, the state plus one questionofficial Qwen weights, frozen, letters readout
DecisioGemma 4 12B37.8 ms (b)$6.92 (e)32,768 tokens per promptofficial Gemma weights, frozen, letters readout
DecisioGemma 4 31B56.0 ms (b)$10.28 (e)32,768 tokens per promptofficial Gemma weights, frozen, FP8 on load, letters readout
Jev 1.13524.1 ms (c)$6.05 (f)64k tokens per request; 32k tokens for the state plus the longest question (g)hosted model, TypeSafe's API
Surogate Rune26B-A4B v3120.5 ms (d)$16.53 (e)not on the boardfull fine-tune of google/gemma-4-26B-A4B-it
Decider chatGemma-4-31B108.5 ms (d)$22.18 (e)not on the boardan inference technique on google/gemma-4-31B-it
pplx-decider-v1-27b101.4 ms (d)$24.94 (e)not on the boardfull fine-tune of Qwen/Qwen3.8-27B

(a) JevBench's board reports Jev's hard tier over 220 questions, 109 of them held out and never published, so it is not the same item set as the 111 published ones we can run.

(b) Our server time on one RTX PRO 6000 Blackwell: the Decision Index kit's HTTP wall time to a server on the same machine, one request at a time, over all 150,759 requests.

(c) The board's own measurement of TypeSafe's hosted API: an HTTPS round trip over the internet from its lab, one request at a time, on its latency sample. It includes the network, so it is not the same measurement as (b) and (d), and it carries no mark.

(d) The board's on-card median on one RTX PRO 6000, one request at a time, on its latency sample.

(e) Self-hosted, estimated as mean request latency × 150,759 requests at $1.48 per card-hour, what the card for our runs cost on vast.ai; for the other entries, with the board's mean latency. One request at a time throughout; a server answering requests concurrently would cost less per pass.

(f) The board's recorded API cost of Jev's run over the same benchmarks, at TypeSafe's tariff of $0.042 per million input tokens, output tokens free (jev.results[].usd).

(g) From TypeSafe's documentation (section 3).

(h) Not on JevBench's board.

Capabilities

Each statement about Jev is what TypeSafe's public documentation states, quoted in snapshots/typesafe_docs.json with each page's hash, read 2026-10-05 at 21:44 UTC. Where a row says the documentation does not state something, that is about the pages read, not a claim about the service.

Capabilities
decisio (all threebases unless noted)Jev 1.13, as TypeSafe'sdocumentation states it
Question typesyes/no (noul), choice and score, on the same /v1/systemone wire format"The three TypeSafe question types (Choice, Score, Noul)" (primitives)
Options per questionup to 255 named options per choice"a maximum of 255 options per Choice"; a Score "accepts up to 10" levels (API)
Several questions per requestthe state is read once and shared by every question in the request"System One models evaluate every question in a request in parallel." (primitives)
Context32,768 tokens per prompt: the state plus one question"64k tokens per request; 32k tokens for state plus the longest question" (models)
Self-hostingyes: one 96 GB card on vLLM, the Docker image, or a Mac (below)served by "the same endpoint, POST /v1/systemone"; the pages read describe the hosted API and no self-hosted option (models)
Open weightsthe bases' official weights, downloaded at pinned revisions, never redistributed"the same weights serve every account"; the pages read name no downloadable weights (models)
Teaching from labelled examplesPOST /v1/tasks fits a per-task calibration and, for long option lists, an intent head on your labelled examples; the model stays frozen"Jev is not fine-tuned or LoRA-adapted with customer data", and "You shape its answers to your domain through the request rather than through per-account weights" (models)
Image inputphotos in the state, through a second engine on the same card (--image-model); measured on the Qwen base only"Text only. String, JSON object, or array of text values. No image, audio, or video input." (models)
Repeatabilitycheckpoints at pinned revisions; Qwen: a repeated request returns the same probabilities bit for bit; Gemma 4 12B: the same choice, probabilities within 0.035; Gemma 4 31B: a repeated request returns the same answers with the engine in the server's process, within one session on one machine; across sessions near-tied answers can move (EVAL_CARD.md sections 5 to 7)"An alias moves when a new release ships, so the answers behind it can change without a change on your side." The pages read do not state whether a repeated request returns the same probabilities (models)
Calibrationfitted temperatures per base; JevBench ECE on the standard and hard tiers: Qwen 0.121 and 0.043, Gemma 4 12B 0.033 and 0.085, Gemma 4 31B 0.035 and 0.091"System One models are trained for calibrated decisions: their probabilities are optimized against outcomes to reflect uncertainty." (System One)
Where data goesto the machine that runs the server, and nowhere else"Jev is not trained on customer requests or responses."; zero data retention is offered "for enterprise customers" (models, legal)
Mac and OllamaQwen and Gemma 4 12B on Apple silicon with MLX (not the 31B); the Qwen base on Ollama (aminroudaki/decisio), with Ollama's own prompt and no calibrationthe pages read describe the hosted API only
LicenceApache-2.0; the weights are Qwen's under Apache-2.0 and Google's Gemma 4 under Apache-2.0 with Google's Gemma Prohibited Use Policya hosted service under TypeSafe's customer agreements, which "govern your use of TypeSafe" (legal)
Priceno per-call price; the card's cost (section 1, e)"$42 / $0.042" per billion / million tokens, "Charged per input token. Output tokens are free." (models)

Every index benchmark

One row per benchmark, grouped by area, with the coverage-adjusted value that feeds the index (raw on the board); ForecastBench shows its Brier loss, where lower is better. n is the number of items behind the interval: the smallest of the benchmark's cases, our scored requests and Jev's cases. Each interval is in runs/2026-10-05_comparison/results.json, and tables.md there is these tables as the script writes them.

Bold: level with or ahead of Jev 1.13, by the comparison's rule; plain: behind.

Knowledge & Reasoning

Knowledge & Reasoning
BenchmarknQwen3.6-35B-A3BGemma 4 12BGemma 4 31BJev 1.13Rune26B-A4B v3Decider chatGemma-4-31Bpplx-decider-v1-27b
GPQA Diamond, accuracy1960.510: behind Jev0.388: behind Jev0.520: behind Jev0.7860.4690.4900.495
GSM8K, accuracy1,3190.544: behind Jev0.599: behind Jev0.805: level with Jev0.7990.7970.8260.674
ChessBench, accuracy5,0000.162: level with Jev0.165: level with Jev0.220: ahead of Jev0.1720.1870.2230.172
MuSR, accuracy7520.581: behind Jev0.581: behind Jev0.640: level with Jev0.6610.6490.6420.618
SATA-Bench, case exact accuracy1,6500.232: behind Jev0.264: level with Jev0.323: ahead of Jev0.2640.3480.2930.299
CRUXEval, accuracy5700.625: behind Jev0.654: behind Jev0.832: ahead of Jev0.7300.7440.7930.749
CLadder, accuracy5,0000.622: behind Jev0.634: behind Jev0.764: ahead of Jev0.7260.7270.7460.745
HLE, accuracy5010.114: behind Jev0.106: behind Jev0.150: behind Jev0.2040.1060.1480.120
MMLU-Pro, accuracy12,0320.613: behind Jev0.550: behind Jev0.694: behind Jev0.8270.6720.6960.645
BBH, accuracy5,5070.694: behind Jev0.691: behind Jev0.762: behind Jev0.9290.8110.7620.781

Language Understanding

Language Understanding
BenchmarknQwen3.6-35B-A3BGemma 4 12BGemma 4 31BJev 1.13Rune26B-A4B v3Decider chatGemma-4-31Bpplx-decider-v1-27b
ContractNLI, macro-F11230.733: level with Jev0.743: level with Jev0.726: level with Jev0.7170.7700.7360.781
ANLI, macro-F13,2000.686: behind Jev0.702: behind Jev0.737: level with Jev0.7480.7350.7320.707
WinoGrande, accuracy1,2670.808: behind Jev0.730: behind Jev0.852: behind Jev0.9190.8600.8370.852
HellaSwag, accuracy10,0420.940: level with Jev0.867: behind Jev0.923: behind Jev0.9450.9240.9200.940
ACOS, per-review F14000.174: behind Jev0.155: behind Jev0.215: behind Jev0.2950.2680.1880.204
FinEntity, macro-F19790.803: behind Jev0.905: ahead of Jev0.931: ahead of Jev0.8700.8840.9090.929
iSarcasmEval, Sarcasm F1 · track A, English1,4000.474: level with Jev0.427: behind Jev0.506: level with Jev0.5050.6040.5110.603
VAST, macro-F13,0060.490: behind Jev0.591: behind Jev0.687: ahead of Jev0.6460.7660.7320.708
NLI4CT, macro-F15,5000.772: behind Jev0.813: behind Jev0.835: level with Jev0.8410.8040.8290.848
RAGTruth, F1 on hallucinated class2,7000.702: behind Jev0.696: behind Jev0.783: level with Jev0.7650.7680.7690.803

Retrieval & Classification

Retrieval & Classification
BenchmarknQwen3.6-35B-A3BGemma 4 12BGemma 4 31BJev 1.13Rune26B-A4B v3Decider chatGemma-4-31Bpplx-decider-v1-27b
BANKING77, macro-F13,0800.746: behind Jev0.729: behind Jev0.784: level with Jev0.7970.8380.7910.790
CLINC150+OOS, macro-F15,5000.822: behind Jev0.872: behind Jev0.901: level with Jev0.8930.8740.9120.878
BRIGHT, nDCG@102200.444: level with Jev0.471: level with Jev0.439: level with Jev0.4750.4630.4360.487
Amazon ESCI, macro-F15,0000.450: behind Jev0.535: level with Jev0.540: level with Jev0.5520.5530.5310.551
PhishNChips phishing decisions, accuracy2,0000.752: ahead of Jev0.813: ahead of Jev0.875: ahead of Jev0.6250.8090.8750.600
HoVer, accuracy4,0000.700: behind Jev0.693: behind Jev0.756: ahead of Jev0.7290.8070.7640.742

Tools & Automation

Tools & Automation
BenchmarknQwen3.6-35B-A3BGemma 4 12BGemma 4 31BJev 1.13Rune26B-A4B v3Decider chatGemma-4-31Bpplx-decider-v1-27b
BFCL, case exact accuracy1,6940.960: level with Jev0.962: level with Jev0.979: ahead of Jev0.9580.9480.9800.976
ToolRet, nDCG@106850.625: level with Jev0.622: level with Jev0.613: level with Jev0.6530.6440.6360.674
API-Bank, accuracy5080.817: behind Jev0.870: level with Jev0.854: level with Jev0.8820.8330.8500.841
Home appliance simulator, case exact accuracy880.398: level with Jev0.443: level with Jev0.614: level with Jev0.5230.4660.6250.739
When2Call, accuracy3,6520.714: behind Jev0.761: behind Jev0.772: behind Jev0.8100.7600.7690.817

Arts & Human Taste

Arts & Human Taste
BenchmarknQwen3.6-35B-A3BGemma 4 12BGemma 4 31BJev 1.13Rune26B-A4B v3Decider chatGemma-4-31Bpplx-decider-v1-27b
BPoMP, accuracy8110.904: level with Jev0.910: level with Jev0.894: level with Jev0.9090.9000.9060.939
Humicroedit, accuracy2,6280.608: level with Jev0.617: level with Jev0.625: level with Jev0.6190.6220.6380.624
POP909-CL, accuracy2,0000.142: behind Jev0.060: behind Jev0.250: ahead of Jev0.1660.6610.2790.378
cfcolor, accuracy5,0000.600: behind Jev0.624: behind Jev0.639: level with Jev0.6440.6260.6260.644
ForecastBench, Brier loss (lower is better)10,1390.228: behind Jev0.202: behind Jev0.207: behind Jev0.1740.2030.2020.195
Habermas Machine, accuracy1,6760.447: level with Jev0.465: level with Jev0.484: level with Jev0.4590.4220.4790.418
New Yorker, accuracy5280.688: level with Jev0.633: behind Jev0.739: level with Jev0.7010.7410.7390.703

Hosted chat models on JevBench's 231 published questions

The JevBench operator's published counts for the general chat models it measured through their providers' APIs, by the board's model IDs (revision v1.4.2.2, public_accuracy × 231); they are not decision models, have no Decision Index value, and carry no marks.

Hosted chat models on JevBench's 231 published questions
Model, as the board names itBoard IDCorrect of 231
GPT-6 Luna, low reasoning effortgpt-6-luna-low229
GPT-6 Luna, default medium reasoning effortgpt-6-luna230
GPT-5.6 Luna, low reasoning effortgpt-5.6-luna225
Gemini 3.1 Flash-Litegemini-3.1-flash-lite201
DeepSeek V4.1 Flash, thinking by defaultdeepseek-flash226
Qwen3.8 27B, on Chutesqwen3.8-27b166

Source: docs/comparison.md Jev's GPQA Diamond is the board's jev.benchmarks['25'].raw (0.7857, 154 of 196 scored); the board file also holds jev.results['25'].score = 0.7828, which its GPQA Diamond column and benchmark page display as 78.3. Field: EVAL_CARD.md

Jev leads all three bases on the knowledge-heavy benchmarks, GPQA Diamond, MMLU-Pro, BBH and HLE, and on WinoGrande, ACOS, When2Call and ForecastBench's Brier loss. On the Qwen and Gemma 4 12B bases it also leads most of language and tools, and their indices are 9.9 and 8.5 points below its own. The Gemma 4 31B base is level with Jev on the index and on three of the five areas, ahead on Retrieval & Classification, and ahead on ten benchmarks, among them CRUXEval, CLadder, VAST, FinEntity, HoVer, BFCL and POP909-CL. All three bases are ahead on PhishNChips, and on JevBench's 231 published questions the 31B answers 213 against Jev's 200, with the other two level at 200. Beyond the scores, the difference is in how each runs: decisio runs on your own card or Mac with the data staying there, takes images and learns a recurring question from labelled examples, while Jev is a hosted API with a longer context per request and, at its tariff, a lower cost for one pass over the suite than our one-request-at-a-time estimates.

Each base on the public harnesses

Accuracy and macro-F1, higher is better.

MeasureQwen3.6-35B-A3B (default)Record: runs/2026-10-04_gemma-base/Gemma 4 12BRecord: runs/2026-10-04_gemma-base/Gemma 4 31BRecord: runs/2026-10-04_gemma-4-31b/
On JevBench's 231 published questions, accuracy: easy / standard / hard1.000 / 0.972 / 0.7391.000 / 0.972 / 0.7391.000 / 1.000 / 0.838
Decision Index BANKING77, macro-F10.7460.7290.785
Decision Index CLINC150+OOS, macro-F10.8220.8710.902
Decision Index GPQA Diamond (196 scored), accuracy0.5100.3780.520
Decision Index MMLU-Pro, accuracy0.6130.5490.694
Intent heads from 10 labelled examples per intent, BANKING77 / CLINC1500.840 / 0.912 (six draws)0.832 / 0.908 (six draws)0.844 / 0.970 (three draws)
Image input, ImajevBench v2.0-lite, the 230 answerable itemsRecord: runs/2026-09-27_image-input/0.791not measurednot measured

It went in under the maintainer's decision, past a pre-registered rule it missed by 0.2 to 1.1 points on three of four benchmarks; EVAL_CARD.md section 7 states the rule, the numbers and the reason.

Server time

One question on a new 300-token state, server time (0.8.1's served defaults, the engine in the server's process; one RTX PRO 6000 at 585 W, AMD Ryzen Threadripper 9960X) Server time in ms. Lower is better.
BaseServer time
Qwen3.6-35B-A3B49.9 ms
Gemma 4 12B53.6 ms
Gemma 4 31B80.7 ms

Qwen3.6-35B-A3B, Gemma 4 12B: runs/2026-10-04_gemma-base/Gemma 4 31B: runs/2026-10-04_gemma-4-31b/

One question on a 1,000-token state from the prefix cache, server time (same) Server time in ms. Lower is better.
BaseServer time
Qwen3.6-35B-A3B20.2 ms
Gemma 4 12B24.8 ms
Gemma 4 31B31.7 ms

Qwen3.6-35B-A3B, Gemma 4 12B: runs/2026-10-04_gemma-base/Gemma 4 31B: runs/2026-10-04_gemma-4-31b/

The latency rows were measured in one session on decisio 0.8.1's served defaults: the engine in the server's process, and on the Gemma bases a single question on a new state registering its boundary first (runs/2026-10-06_latency-585w/).

The card was one RTX PRO 6000 Blackwell Workstation Edition with its power limit at 585 W (default 600 W), on an AMD Ryzen Threadripper 9960X host.

On the Gemma bases that limit held the clock back during most first reads, so a card at 600 W may read new states faster; on the Qwen base it did not.

In the same session, the boundary registration added +16 to +28 ms (12B) and +21 to +34 ms (31B) to the first read of a new state of 300 to 3,000 tokens, and a second, different question then read the state from the cache on 20 of 20 states per base.

An earlier change, paired within its own session:

Gemma 4 12B against Qwen3.6-35B-A3B, paired

What drives the choice, measured on one RTX PRO 6000 Blackwell in one session, each base with its own defaults, paired over the same items (95% bootstrap intervals; runs/2026-10-04_gemma-base/):

Gemma 4 12B against Qwen3.6-35B-A3B on the same items
MeasureGemma 4 12BQwen3.6-35B-A3BGemma minus Qwen
JevBench v1.5 open-set reading, I_open (equal types)64.449.4+15.0 [+7.0, +23.6]
JevBench v1.5, yes/no / score49.0 / 63.713.7 / 57.2
Yes/no answers between 0.20 and 0.80 (74 items)15%39%
JevBench, published items, correct0.0 points [-4.3, +4.3]
Decision Index accuracy, CLINC150+OOS+4.5 [+3.6, +5.5]
Decision Index accuracy, BANKING77-1.4 [-2.7, -0.1]
Decision Index accuracy, GPQA Diamond-13.6 [-21.7, -5.6]
Decision Index accuracy, MMLU-Pro-6.4 [-7.2, -5.5]
1,400-item suite, accuracy0.7350.770-3.5 [-5.5, -1.6]
One question on a new 3,000-token state, server time (0.8.1's served defaults, the engine in the server's process; one RTX PRO 6000 at 585 W, AMD Ryzen Threadripper 9960X; one session, runs/2026-10-06_latency-585w/)258.6 ms85.2 ms

Record: runs/2026-10-04_gemma-base/

Calibration

The calibration of each base is in the capabilities table above.

EVAL_CARD.md has the full tables, the calibration figures and the disclosures of what was fitted on what (sections 4, 6.4 and 7.4).

The Decision Index, whole suite: self-runs

Self-run, submitted to the board and pending the maintainers' validation; no rank is claimed.

The Decision Index 0.2.1, whole suite, self-run per base
BaseTagDecision Index 0.2.1Raw indexBreadthRequest latency, median / p95
Qwen3.6-35B-A3B (FP8)v0.4.048.0660.4546.5570.5 / 261.2 ms
Gemma 4 12B (bf16)v0.4.049.4361.4747.5737.8 / 275.9 ms
Gemma 4 31B (FP8 on load)v0.6.057.5867.3356.5056.0 / 439.3 ms

Source: EVAL_CARD.md

What these numbers do not show

Source: EVAL_CARD.md

  • One card per run and one run per configuration.
  • JevBench's v1.5 reading is taken on JevBench's 231 published questions, of the board's 904 open items, with no judge tier and no sealed half; its published items have been public since v1.2 and may be in any model's pretraining data.
  • The Decision Index rows of sections 3, 6 and 7 were rebuilt with the kit for four benchmarks; their counts and the rows' sha256 match (decisio.bench.di_rows), but a partial rebuild cannot be checked against the whole suite's hash. Section 8's runs used the whole suite, rebuilt with the kit and matched against the lab's hashes.
  • The ImajevBench record was made without the rendering rules and the key-order tie-break; 2 of its 254 items ended in exact ties that the harness rejected and are counted wrong.
  • The second-engine head mode with the registered text-only class is unmeasured for latency.
  • Latency is server-side on the card's localhost; cost is the card-hour price divided by measured throughput, with nothing else counted.
  • Latency depends on the host and the card's power limit. The latency rows marked 585 W come from one RTX PRO 6000 Blackwell Workstation Edition whose power limit was 585 W (default 600 W), on an AMD Ryzen Threadripper 9960X host (runs/2026-10-06_latency-585w/). On the Gemma bases the limit held the clock back on 72 to 92% of the busy samples, one request at a time included, so a card at 600 W may read new states faster; on the Qwen base it did on 0 to 2%. A cached Qwen question depends on the CPU instead: 39.4 ms in-process on an AMD EPYC 7452 host, 18.5 ms on an AMD Ryzen 9 9950X and 20.2 ms on the Threadripper, the same card model and vLLM (runs/2026-10-05_engine-death-gates/, runs/2026-10-06_latency-0.8.1/). Latency rows not marked come from hosts whose CPU and power limit were not recorded.
  • The per-item Decision Index records keep each request's id, the payload's sha256 and our response, not the item text (GPQA's authors ask that its items not be published in plain text).

Every table

The evaluation card has the full tables for each base, the calibration figures and what was fitted on what: the evaluation card.