Mac with MLX
uv sync --extra mlxuv run python -m decisio.serve.vllm_engine --backend mlx --model mlx-community/Qwen3.6-35B-A3B-6bituv run python -m decisio.serve.vllm_engine --backend mlx --base gemma-4-12b --model mlx-community/gemma-4-12B-it-6bitNeeds a Mac with Apple silicon: taken verbatim from the repository's tested docs, run record runs/2026-10-03_mlx-regate/.
--backend mlxserves the text route on Apple silicon from an MLX conversion of either base's checkpoint, with every feature of that route: the letters readout, the shared state prefix, the temperatures, task registration with calibration and the intent head, abstention, the rendering rules and the tie-break.- The base comes from
--base, or from the conversion's model type, and brings its own settings, as on vLLM;--modelnames the conversion. - It refuses the image route, packed mode, LoRA adapters and the second-engine head.
The Qwen base:
- The 6-bit conversion is the Mac default and the 4-bit one the option for 32 GB machines; the 8-bit one is not shipped.
- Peak memory at 32,761 tokens of state: 30.7 GB at 6 bits, 22.0 GB at 4 bits, 39.4 GB at 8 bits (
runs/2026-10-02_mlx-backend/). - It does not pad a state (
--pad-policy none, the MLX default; its cache needs no padding), so its prompts are the served ones without the padding, and a question asked alone and inside a request is one prompt. - A cross-request prefix cache (
--prefix-cache-mb, 2048 by default; 0 turns it off) lets a request whose state was seen before continue from the kept cache, with the same answers bit for bit. - The prompts are built with the official tokenizer by default (
--tokenizer), so they are the vLLM path's byte for byte. - At 6 bits, under the current served default, it passes the gates against the FP8 records except the pooled-ECE gate, which changes no setting (
runs/2026-10-03_mlx-regate/summary.mdexplains why). - Server time on an Apple M5 Pro (64 GB), median: 266 ms for one question on a new state, 111 ms on a state seen before, 130 ms per question when 100 share a state (
runs/2026-10-03_mlx-regate/).
The Gemma base (--base gemma-4-12b):
- The 6-bit conversion is the default; there is no 4-bit option, since the 4-bit conversion fails the suite-accuracy gate (
runs/2026-10-05_mlx-gemma/). - The 6-bit conversion holds 9.7 GB of weights and peaks at 15.1 GB at 32,687 tokens of state, so it runs on a 32 GB Mac (
runs/2026-10-05_mlx-gemma/6bit/memory.json). - The prompts are built with the base's own tokenizer and chat template,
google/gemma-4-12B-itat its pinned revision, not the conversion's, whose chat template is older. - At 6 bits it passes every gate against the base's vLLM bf16 record: suite accuracy and ECE, the intent heads over six draws, conformance, and answers bit for bit on a fresh server (
runs/2026-10-05_mlx-gemma/summary.md). - It reads a new state more slowly than the Qwen base on the same Mac: one question on a new 1,000-token state takes 1,299 ms, against 704 ms for the Qwen base (server medians on an Apple M5 Pro under load;
runs/2026-10-05_mlx-gemma/,runs/2026-10-03_mlx-regate/).
docs/design/mlx-backend.md has the design, the gates of each conversion and the padding decision.
Decisio is an independent open-source project, not affiliated with or endorsed by TypeSafe. Jev is TypeSafe's product name.
Apache-2.0. This page is built from Decisio v0.8.2.