Questions
What does it do?
Section titled “What does it do?”Each question becomes one prompt, the state followed by the question and its lettered options, and its answer is the model's distribution over the option letters at one position, so every question costs one forward pass and no generated text.
The state is prefilled once per request and every question reads it from vLLM's prefix cache.
decisio's vLLM plugin registers its model classes through an entry point, so vLLM loads them without a patch, and the served class also returns the hidden state at the answer position for the intent head.
The design notes: docs/design/vllm-plugin.md, docs/design/hidden-state-readout.md and docs/design/mlx-backend.md.
Is it affiliated with TypeSafe?
Section titled “Is it affiliated with TypeSafe?”No. It is an independent open-source project, not affiliated with or endorsed by TypeSafe, and Jev is TypeSafe’s product name. It serves TypeSafe’s published System One wire format, so a System One client can be pointed at it: see Migrating a System One client.
Which base should I choose?
Section titled “Which base should I choose?”Choosing a base compares the three, with each base’s provenance and what drives the choice.
What hardware do I need?
Section titled “What hardware do I need?”| Machine | Path | Memory needed |
|---|---|---|
| Linux with one NVIDIA card | GPU server | a 96 GB card, as measured; each base's weights are in Choosing a base |
| Linux with one NVIDIA card and Docker | Docker | as the GPU server |
| Mac with Apple silicon | Mac with MLX | each base's, in docs/running.md |
| Mac, Linux or Windows with Ollama | Ollama | the size of the tag, on the model page |
| Any machine, for development | CPU stand-in | a small Hugging Face model on the CPU |
docs/running.md has the details of every path.
Can it reason, or answer knowledge questions as a chat model does?
Section titled “Can it reason, or answer knowledge questions as a chat model does?”Each question is one forward pass with no text generated, so nothing is worked out step by step. How it compares shows what that costs on hard knowledge questions, and Pong’s limit shows it in a game.
Are the benchmark figures board scores?
Section titled “Are the benchmark figures board scores?”No: they are self-runs with the public harnesses, not board scores. Benchmarks names the record behind every figure, and the evaluation card says what was fitted on what.
Does it learn from my requests?
Section titled “Does it learn from my requests?”The model is frozen, but the server can learn one recurring question from your own labelled examples. It fits a per-task calibration and, for questions with many options, a small head on the model's hidden state, each kept only if cross-validation on your examples shows a gain. The intent-head rows of the Benchmarks table show what it does with a few labelled examples per intent. Tasks registered under the earlier compact layout are not applied by the current default and need registering again.
curl http://127.0.0.1:8000/v1/tasks -H 'Content-Type: application/json' -d @examples/tasks/examples.jsonRun against the CPU stand-in at v0.8.2 on 2026-10-07: registers the walk-through task again.
docs/tasks.md walks through it and states what each number was measured on; examples/tasks/ runs it on a laptop against the CPU stand-in.
Do repeated requests return the same answer?
Section titled “Do repeated requests return the same answer?”On a running server with the Qwen base, a question returns the same probabilities every time, alone or among other questions: each question is scored in its own engine call.
Across server restarts, single-question answers have matched an earlier record on the whole suite in most starts measured, and moved slightly without changing a choice in the others.
The Gemma base does not repeat bit for bit: its logits come out in bf16 after its soft cap, and a state read for the first time and read again from the prefix cache can land a step apart; several questions in one request stay within the bound in the README's settings table.
Any change of rounding moves this model's answers, the batch, the padding or where a long row is split, so every bit-identity claim holds for a fixed configuration and request form.
Measurements: EVAL_CARD.md sections 4 and 6.5.
Where does my data go?
Section titled “Where does my data go?”The server runs where you start it and listens on 127.0.0.1:8000 unless --host and --port say otherwise.
Decisio is a model server, not a hardened public endpoint.
- It is designed to run behind your own proxy on a private network. The server binds to
127.0.0.1by default; do not expose it directly to the internet. - There is no authentication on any route. Put authentication, rate limiting and request-size limits in the proxy.
POST /v1/tasks,POST /v1/tasks/import,DELETE /v1/tasks/{id}and the abstention routes change server state. Restrict them to trusted callers.--debug-readoutexposes raw model readouts; leave it off in production.- The image route accepts image payloads from requests; size limits belong in the proxy.
- vLLM's own security guidance applies to the engine underneath: https://docs.vllm.ai/en/stable/usage/security/
A decisio server is single-tenant. Its prefix cache and registered tasks are shared by every caller, so response timing can reveal whether another caller recently sent the same text, and one caller's registered task answers another's identical question. Run one server per trust boundary.
How do I report a security problem?
Section titled “How do I report a security problem?”Report security issues through GitHub's private vulnerability reporting on this repository ("Report a vulnerability" under the Security tab). Do not open a public issue.
You will get an acknowledgement within 5 working days and an assessment within 15. Fixes for confirmed issues are released as soon as they are ready, with credit to the reporter unless you ask otherwise.
What is the licence?
Section titled “What is the licence?”Apache-2.0 (LICENSE, NOTICE).
The model weights are Alibaba's Qwen3.6-35B-A3B under Apache-2.0 and, for the other two bases, Google's Gemma 4 12B and 31B under Apache-2.0 with Google's Gemma Prohibited Use Policy; all are downloaded, not redistributed.
THIRD-PARTY.md lists everything else this project builds on.
How can I contribute?
Section titled “How can I contribute?”See CONTRIBUTING.md.
Pull requests need a DCO sign-off (git commit -s), pass the CPU tests and lint, and are reviewed by a maintainer before merge.
Security issues go through GitHub's private vulnerability reporting, as described in SECURITY.md.
Decisio is an independent open-source project, not affiliated with or endorsed by TypeSafe. Jev is TypeSafe's product name.
Apache-2.0. This page is built from Decisio v0.8.2.