GPU server
Requirements: Linux, one NVIDIA card (measured on an RTX PRO 6000 Blackwell with 96 GB), a driver that supports CUDA 13.0 (the runtime uv.lock pins), Python 3.12 and uv.
git clone https://github.com/aminry/decisiocd decisiouv sync --extra serve --frozenuv run python -m decisio.serve.vllm_engine --base qwen3.6-35b-a3bNeeds Linux with one NVIDIA card: taken verbatim from the repository's tested docs, run record runs/2026-10-01_quickstart-public/.
--base qwen3.6-35b-a3bservesQwen/Qwen3.6-35B-A3B-FP8and--base gemma-4-12bservesgoogle/gemma-4-12B-it, each at its pinned revision and with its own settings (the README's "Choosing a base";docs/cli.mdlists every setting per base).- Without
--base, the base is detected from--model'sconfig.json, so a local copy of either checkpoint brings its own settings; a base's pinned revision applies whenever--modelnames the base's own Hugging Face repository and--revisionis not given. - The first start downloads the checkpoint (about 36 GB for the Qwen base,
runs/2026-10-01_docker-first-gpu-start/) into the Hugging Face cache and warms the engine;/healthanswers once it is ready. - The decisio plugin registers its model classes with vLLM through an entry point, so vLLM 0.30.0 loads them without any patch; the package must be installed (as
uv syncdoes) for vLLM's engine processes to find it (docs/design/vllm-plugin.md). patches/holds an optional latency patch series, off by default and not needed for correct answers (patches/README.md).- The server refuses to start when
VLLM_USE_DEEP_GEMMis set to anything other than0, unless--allow-deep-gemmis given. - It listens on
127.0.0.1:8000(--host,--port); put a reverse proxy in front of it to expose it. --image-modeladds a second engine on the same card for requests that carry images;EVAL_CARD.mdsection 1 has the memory shares it was measured with.- vLLM's engine runs in the server's process by default; with a second engine (
--image-model,--head-engine,--one-engine) it runs in a process of its own (--engine-process,docs/cli.md). - A later, different question about a state read before, sent in its own request, reads the state from the prefix cache. On the Gemma bases vLLM keeps only a finished request's latest sliding-window checkpoint, which lies inside its question, so a single question first registers its state's boundary with the warm-up that multi-question requests already send (the state and one token). The server remembers which states it registered and sends the warm-up again only when a request finds the boundary gone, so a repeated question costs no extra engine call. 0.8.1 made that second question read the state from the cache: on one card at 400 W with an AMD Ryzen 9 9950X, 29.2 ms on Gemma 4 12B and 37.1 ms on Gemma 4 31B after first reads of 348 and 615 ms on a 3,000-token state, where before it read the whole state again (531 ms instead of 47 at 3,000 tokens and 9.2 s instead of 0.17 at 31,000 on the 31B, on another host). In exchange the registration adds to the first read of a new state: +11, +37 and +59 ms at 300, 1,000 and 3,000 tokens on the 12B, and +23, +43 and +72 ms on the 31B, on the same host (
runs/2026-10-06_latency-0.8.1/). A repeated question and several questions in one request were not affected. - vLLM's
prefix_cache_retention_intervalwould also keep the boundary, but it keeps every sliding-window block of each state, and with the cache filled to 1.3 times its pool the earliest states lost even their repeated question, where the default (0) kept them at twice the pool; it stays at the default. The Qwen base needs neither: its state is padded to end on its 1,056-token block, the latest checkpoint of a request with one question./healthreportscache_hit_unit,hash_unitandregisters_state_boundary.
Per-request latency depends on the host's CPU and on the card's power limit, as well as the card.
- The CPU: with the engine in the server's process, a Qwen question on a cached state took 39.4 ms of server time on an AMD EPYC 7452 (Zen 2) host and 18.5 ms on an AMD Ryzen 9 9950X (Zen 5) host, with the same card model and vLLM (
runs/2026-10-05_engine-death-gates/,runs/2026-10-06_latency-0.8.1/). The Gemma bases differed far less between the two hosts (Gemma 4 12B: 28.0 against 23.4 ms). - The power limit: first reads are bound by the card. On the Ryzen host the card was limited to 400 W of its 600 W, and its first reads on the Gemma bases were slower than earlier records whose power limit was not captured.
Quote latency with the host's CPU model and the card's power limit beside it (nvidia-smi -q -d POWER).
Decisio is an independent open-source project, not affiliated with or endorsed by TypeSafe. Jev is TypeSafe's product name.
Apache-2.0. This page is built from Decisio v0.8.2.