Skip to content

Deploying and operating

curl http://127.0.0.1:8000/health returns what the server is serving: the engine, the base, the checkpoint and its revision, the temperature for each question type and the prompt. docs/api.md describes every field.

What the server is serving, in one object a run record can quote:

  • ok, and the engine's facts: the vLLM version, the GPU, the mode, the padding (pad_unit, pad_where), the quantization and the cache data types.
  • block_size and match_unit, the KV cache's block and the unit a prefix match is counted in; cache_hit_unit, the step prefix-cache hits come in, and hash_unit, the step prefixes are hashed at (on the Gemma 4 12B base, whose cache keeps groups of 16- and 64-token blocks, hits come in 64-token steps; 32 on the 31B; 1,056 on the Qwen base); registers_state_boundary, whether a single question registers its state's boundary in the cache for the next question (docs/running.md).
  • base and profile: the base, the checkpoint and its revision, where vLLM's engine runs (engine_process), the temperature for each question type, and the prompt (layout, answer position, label forms, system turn, yes/no and option rendering, multi-question scoring, padding).
  • prompt_format: the prompt's layout settings.
  • systemone: the System One route's settings (rendering rules, noul_commit, abstention and its tasks, orders, tasks and which are applied, the temperatures).
  • head_engine, the intent head's hidden-state reader (the serving engine by default, a second engine with --head-engine), and image_engine when one is running.

Once the engine has died, /health answers 503 with {"ok": false, "engine": "dead", "reason": "...", "exit_code": 70}, every other route answers 503 with the reason, and the server exits with code 70 (docs/running.md, "When the engine dies").

An engine whose forward pass fails (a CUDA error, an out-of-memory error) or whose engine-core process exits does not serve again, so the server ends itself and leaves the restart to whatever runs it.

A death is confirmed before it is declared, so that a request which trips a bug cannot take a healthy server down:

  • vLLM's own word is enough: its EngineDeadError, or its flag for an engine-core process that is gone.
  • Any other unexpected error fails that request with 500, as it always has, and the server then sends the engine one probe: a forward pass over a one-token prompt, given 5 seconds, or three times the engine's slowest recent call when that is longer, so that a slow engine is not taken for a dead one (the server times one probe per engine at start-up, so its first request already allows for a slow engine).
  • If the probe answers, the engine is alive and the server stays up; the server log says so (the engine answered a probe in ... ms).
  • If the probe fails or does not answer in that time, the engine is dead. An in-process engine that has failed hangs rather than answering, so this is how its death shows.
  • A probe still waiting when the engine is declared dead another way (vLLM's flag, another request) stops at once, and its request gets its 503.
  • An error caused by the request (a malformed question, a state longer than the context) is answered with a 4xx and the engine is not probed.

From the moment the engine is dead:

  • every request, including those already waiting for the engine, is answered at once with 503 and the reason, and the engine is not called again;
  • /health answers 503 with {"ok": false, "engine": "dead", "reason": "...", "exit_code": 70};
  • after 2 seconds, once every request in flight has its answer (waiting at most 8 seconds more), the server stops its engine processes and exits with code 70; a request the dead engine never returns has its connection closed by the exit.

The in-process arrangement therefore exits about 7 seconds after the failing request (the probe's 5 seconds, then the grace period's 2). In the separate arrangement (--engine-process separate, the default with a second engine) vLLM reports the death itself, so the server exits about 2 seconds after it. The server also watches vLLM's flag between requests, so a death between requests turns /health to 503 within a second, before any request arrives.

Run the server under a restart policy, so that the exit brings up a new one:

  • compose: restart: unless-stopped, as compose.yaml has it;
  • systemd: Restart=on-failure;
  • Kubernetes: the default restartPolicy: Always, with a liveness probe on /health.

Docker restarts a container when its process exits, not when its HEALTHCHECK reports it unhealthy, so it is the exit that brings the server back; the image's check reports the dying container unhealthy in the meantime. The new server loads the engine again (the start times above), and requests sent until it is healthy are refused.

In decisio 0.1.0 to 0.7.1 none of this happens: after the engine dies /health keeps answering 200 and the process keeps running, so no restart policy acts. In the separate arrangement every request then fails with 500; with --engine-process in (0.7.0 and 0.7.1) every request after the failing one hangs.

Registered tasks live in the server’s memory. Teaching it a task shows how to export them and load them again at start.

Report security issues through GitHub's private vulnerability reporting on this repository ("Report a vulnerability" under the Security tab). Do not open a public issue.

You will get an acknowledgement within 5 working days and an assessment within 15. Fixes for confirmed issues are released as soon as they are ready, with credit to the reporter unless you ask otherwise.

The latest release on main receives security fixes. Older tags do not.

Decisio is a model server, not a hardened public endpoint.

  • It is designed to run behind your own proxy on a private network. The server binds to 127.0.0.1 by default; do not expose it directly to the internet.
  • There is no authentication on any route. Put authentication, rate limiting and request-size limits in the proxy.
  • POST /v1/tasks, POST /v1/tasks/import, DELETE /v1/tasks/{id} and the abstention routes change server state. Restrict them to trusted callers.
  • --debug-readout exposes raw model readouts; leave it off in production.
  • The image route accepts image payloads from requests; size limits belong in the proxy.
  • vLLM's own security guidance applies to the engine underneath: https://docs.vllm.ai/en/stable/usage/security/

A decisio server is single-tenant. Its prefix cache and registered tasks are shared by every caller, so response timing can reveal whether another caller recently sent the same text, and one caller's registered task answers another's identical question. Run one server per trust boundary.

Decisio is an independent open-source project, not affiliated with or endorsed by TypeSafe. Jev is TypeSafe's product name.

Apache-2.0. This page is built from Decisio v0.8.2.