Skip to content

Flags and options

The served defaults need no flags; these change behaviour (docs/cli.md has every flag and its default per base):

  • --pad-policy row: pads a single-question request so its whole row ends on the block boundary, faster on a new state at a small cost in repeatability.
  • --multi-question warm: scores a request's questions in one batch for bulk scoring, faster, but each answer then depends on the batch. It is the Gemma 4 31B base's default: with the engine in the server's process a warm request repeats exactly within one session on one machine (not across sessions, EVAL_CARD.md section 7.5), and on the items measured it chose as sequential did every time (runs/2026-10-05_engine-death-gates/).
  • --noul-commit: reports a yes/no probability inside JevBench v1.5's no-answer band at the band's edge; the answer never changes, its calibration does.
  • --prompt-tail compact: the earlier layout, for tasks registered under it.
  • --image-model: a second engine on the same card for requests that carry images.
  • --backend: vllm (the default), mlx for Apple silicon, hf for the CPU stand-in.
uv run python -m decisio.serve.vllm_engine [flags]

A shape to fill in, not a command to run.

--base chooses the base model and brings its checkpoint at a pinned revision and its value for every setting marked "the base's" below; a flag given explicitly overrides it. Without --base, the base is detected from --model's config.json; with neither, the server refuses to start. The served defaults need no other flag: every default below is what was measured, base by base (EVAL_CARD.md sections 1, 6.1 and 7.1). gemma-4-31b is quantized to FP8 when it loads (vLLM 0.30.0's FP8 on load), a setting of its profile; --engine '{"quantization": null}' loads it at bf16, which does not leave room for a 32,768-token context on a 96 GB card (EVAL_CARD.md section 7). --help prints the same list with each flag's description.

Where the columns differ, the value is the base's.

Flagqwen3.6-35b-a3bgemma-4-12bgemma-4-31bWhat it does
--basethe base; without it, detected from --model
--modelQwen/Qwen3.6-35B-A3B-FP8google/gemma-4-12B-itgoogle/gemma-4-31B-itthe checkpoint, a local directory or a Hugging Face repository id
--revision95a723d08a9490559dae23d0cff1d9466213d989707f0a3b8a3c7ad586ed01e27eafbad8a27dd0f7842da3794eaa0b77d5f08bae87a17459d91ff475the checkpoint's revision on the Hub; the pin applies whenever --model is the base's own repository
--backendvllmvllmvllmvllm; mlx for Apple silicon (the Qwen base and gemma-4-12b, docs/running.md; not gemma-4-31b); hf for the CPU stand-in, not for measurement
--model-classhidden-readouthidden-readouthidden-readouthidden-readout: decisio's text class that also returns the hidden state at the answer position, for the intent head; text-only: the same without it (the default with --head-engine); view: a directory built by decisio.serve.make_text_only, loaded as it is, needing no plugin
--modeseparateseparateseparateseparate: one prompt per question, the state shared through the prefix cache; packed: questions packed into one pooling request (the compact layout only)
--pack161616questions per pack in packed mode
--engine{"compilation_config": {"max_cudagraph_capture_size": 4096}}the samethe sameextra LLM(...) keyword arguments, as JSON
--adapternonenonenonename=path of a vLLM-format LoRA adapter; repeatable
--gpu-memory-utilization0.90.90.9the text engine's share of the card
--allow-deep-gemmoffoffoffstart even when VLLM_USE_DEEP_GEMM is set to something other than 0
--served-namedecisio-qwen3.6-35b-a3b-lettersdecisio-gemma-4-12b-it-lettersdecisio-gemma-4-31b-it-lettersthe name GET /v1/models lists
--host, --port127.0.0.1, 8000the samethe samethe listen address; put a reverse proxy in front to expose it
Flagqwen3.6-35b-a3bgemma-4-12bgemma-4-31bWhat it does
--prompt-tailspacedspacedspacedspaced: a blank line before and after the lettered options, then one line asking for the chosen option's letter alone; compact: the earlier layout, for tasks registered under it
--answer-slotprefilltemplatetemplatewhere the label is read: after "Answer:" prefilled in the assistant turn, or at the chat template's own first assistant position
--label-variantssinglesummedsummedthe tokens read per label: the one form the slot reads, or every single-token form summed
--system-promptoffonona system turn before each question
--noul-renderingletters-keysletterslettershow a yes/no question is asked: letters-keys, a two-option letter choice, the false side first, the sides named; letters, the same with each side shown as its description; words, the earlier rendering, read from the yes and no tokens
--describe-optionsonononshow an option that has a description as its description alone; --no-describe-options shows key: description
--hide-index-keysonononnever show enumerated keys (option_0, option_1, ...)
--desnake-labelsonononshow bare snake_case labels as words
Flagqwen3.6-35b-a3bgemma-4-12bgemma-4-31bWhat it does
--temperature1.5063.5925.252the temperature on the text route's plain readout, softmax(log p / T); 1 switches it off; never changes the most probable option
--temperature-choice1.370the global one4.672the temperature for choice questions
--temperature-noulthe global onethe global onethe global onethe temperature for yes/no questions
--temperature-scorethe global onethe global onethe global onethe temperature for score questions
--noul-commitoffoffoffreport an uncommitted yes/no probability at the band's edge (below)
--orders1112: two-order averaging, each question also read in a second option order
--branch-lognonenonenonea JSONL file for the two orders' disagreement per question
Flagqwen3.6-35b-a3bgemma-4-12bgemma-4-31bWhat it does
--multi-questionsequentialsequentialwarm with the engine in the server's process, else sequentialhow a request's questions are scored (below)
--engine-processin (separate with a second engine)in (separate with a second engine)inwhere vLLM's engine runs: in, the server's process; separate, a process of its own (vLLM's arrangement) (below)
--pad-policyalwaysalwaysalwaysalways: pad every state; shared: pad only requests with more than one question; row: pad a single-question request's whole row (below); none: never pad (the default with --backend mlx)
--pad-toblocknonenonewhat a state is padded to: the KV cache block, a token count, or nothing
--pad-wherefrontfrontfrontwhere the padding goes: front, before the chat template; user, at the start of the user turn; between, between the state and the question
Flagqwen3.6-35b-a3bgemma-4-12bgemma-4-31bWhat it does
--tasksonononapply registered per-task calibration and intent heads (POST /v1/tasks)
--tasks-filenonenonenonea JSON list of tasks to load at start, as GET /v1/tasks?full=1 returns them
--head-engineoffoffoffread the intent head's hidden state from a second copy of the model in vLLM's pooling mode instead of from the serving engine: faster head questions, at a second weight copy; not with --image-model
--head-gpu-memory-utilization0.470.470.47the second engine's share of the card
--abstentiononononapply registered abstention thresholds (POST /v1/abstention/tasks)
--abstention-tasksnonenonenonea JSON list of abstention tasks to load at start
--abstain-optionnonenonenoneoffer this extra option (for example "can't tell") on requests that use imajev's extension, reported as unknown_probability and abstained; untrained
Flagqwen3.6-35b-a3bgemma-4-12bgemma-4-31bWhat it does
--image-modelnonenonenonethe full multimodal checkpoint, as a second engine serving only requests that carry images; measured with the Qwen base only
--image-gpu-memory-utilization0.50.50.5the image engine's share of the card
--one-engineoffoffoffload only the image engine and serve text requests on it too, for a card that cannot hold both; changes the text route's class
FlagDefaultWhat it does
--prefix-cache-mb2048the cross-request prefix cache's budget in MB; 0 turns it off
--tokenizerQwen/Qwen3.6-35B-A3B-FP8the tokenizer the prompts are built with, so they are the vLLM path's byte for byte
FlagDefaultWhat it does
--debug-readoutoffhonour the x-decisio-debug request header (the raw readout in the response); never in production

Several questions in one request (--multi-question)

Section titled “Several questions in one request (--multi-question)”

With sequential, the served default, a request with several questions first prefills the state once, then scores each question in its own engine call, reading the state from the prefix cache, so every answer equals the same question sent alone. With warm, for bulk scoring, the questions are scored in one batch after the same prefill: faster (four questions on a new state take about half the time of sequential on the Qwen base and about two thirds on the Gemma bases), but a batched answer differs from the same question sent alone, and when two options are close the choice can change (EVAL_CARD.md section 4). With the engine in its own process (--engine-process separate) a warm answer can also move between repeats, because that process does not always put a request's questions in one engine step (vllm-project/vllm#59764); with the engine in the server's process (--engine-process in) an identical request repeats exactly. --engine-process resolves to in for a single-engine server (since 0.8.0) and to separate when a second engine is configured (--image-model, --head-engine, --one-engine), with which in is refused; an explicit choice is honoured, and the server prints the arrangement and why at start-up and reports it in /health as engine_process. On the Gemma 4 31B base warm is the default when the engine runs in the server's process: against sequential it changed no choice on the suite, JevBench and the travel requests, by at most 0.0073 in a probability (runs/2026-10-05_engine-death-gates/); the other bases keep sequential. With batch, every question goes in one engine call that prefills the state itself, with no warm-up. Measurements: EVAL_CARD.md section 4, runs/2026-10-02_multi-question-and-rendering/, and the README's example request in runs/2026-10-01_docker-first-gpu-start/repeat_variability/.

By default every state is padded to end on the KV cache block, so a single question on a state the server has not seen is prefilled in two engine steps, the state and then the question. With row, a single-question request is padded so that its whole row, state and question, ends on the block boundary, and vLLM prefills it in one engine step; requests with several questions are padded as by default. It is faster on a new state, and accuracy and calibration stayed within noise on the suite, JevBench, the four Decision Index benchmarks and the intent heads. Its cost is repeatability: a repeated identical request can move slightly on long prompts, where the default padding returns the same probabilities every time, so it is off by default. The demos ran with it (docs/demos/README.md). Measurements: runs/2026-10-02_pad-policy-row/; not measured on the Gemma base.

On a running server with the Qwen base, a question returns the same probabilities every time, alone or among other questions: each question is scored in its own engine call. Across server restarts, single-question answers have matched an earlier record on the whole suite in most starts measured, and moved slightly without changing a choice in the others. The Gemma base does not repeat bit for bit: its logits come out in bf16 after its soft cap, and a state read for the first time and read again from the prefix cache can land a step apart; several questions in one request stay within the bound in the README's settings table. Any change of rounding moves this model's answers, the batch, the padding or where a long row is split, so every bit-identity claim holds for a fixed configuration and request form. Measurements: EVAL_CARD.md sections 4 and 6.5.

For scorers that treat a yes/no probability strictly between 0.20 and 0.80 as no answer, as JevBench v1.5 does, such an answer is reported at the band's edge on its own side: 0.80 above 0.5, and 0.20 at or below it. The answer never changes; its probability does, and calibration pays for it: log loss and ECE rise, while JevBench v1.5's yes/no competence rises. Measurements: runs/2026-10-03_noul-commit/ and docs/handoffs/tasks.md.

The earlier layout (--prompt-tail compact)

Section titled “The earlier layout (--prompt-tail compact)”

The compact layout has no blank lines around the options and ends with "Answer with the letter only.". Tasks registered under one layout are not applied under the other, so a task registered before the spaced layout became the default needs --prompt-tail compact or registering again. Packed mode reads the compact layout only, and the image route keeps it.

Decisio is an independent open-source project, not affiliated with or endorsed by TypeSafe. Jev is TypeSafe's product name.

Apache-2.0. This page is built from Decisio v0.8.2.