A typed answer is not permission to act

On October 2, 2026, ggml-org published New in llama.cpp: Decision Models. It describes a local server endpoint, /v1/systemone, where a caller supplies a state and typed questions and receives probabilities over supplied options. The implementation trail is concrete: pull request #29818 was opened October 1, merged October 2, and produced commit a4cb4c6. The PR describes decision-model support as small, self-contained handling around embedding-style models and GGUF decision metadata.

This is a local runtime integration boundary, distinct from the earlier JSON-schema parser update. JSON Schema constrains generated syntax. /v1/systemone returns a choice, score, or yes/no-shaped judgement with model probabilities. Neither establishes that the judgement is true, current, calibrated for a business, or authorized to trigger a side effect.

  1. 1Observed facts and deterministic policy
  2. 2fixed candidate options
  3. 3one typed question
  1. 1Pinned local decision model
  2. 2probabilities and typed answer
  3. 3code checks threshold and authority
  1. 1Low confidence, unsupported model, invalid request, or stale policy
  2. 2unavailable result
  3. 3no mutation
Consider the sequence and each role.

Read the server contract before using a model

The current server README is an operating contract, not a dated release. It says questions are independent except for Clef's joint evaluation. Pin the model and revision, then inspect that model's option ceiling and truncation behavior before accepting a request; a shared endpoint is not an interchangeable classifier. A choice carries the highest-probability supplied option, distribution, and confidence; score is probability-weighted and noul is a true-probability. output_tokens is always zero, while stored temperature does not guarantee calibration on caller data. Invalid requests are 400; a non-decision model, or unsupported image input, is 501. Record these as unavailable contract states. Do not silently select text generation, another model, or a remote provider.

The llama.cpp repository license is MIT. It does not transfer to model weights. The inspected visible OpenJev GGUF card displays CC-BY-NC-4.0; its page creation time is not treated as a publication date. Check the specific card, revision, terms, and intended use before a download or deployment. This article did not download a model, accept a model license, start an endpoint, or make an API call.

An unexecuted N=1 offline fixture

This is an original worksheet, not a measured result or a server procedure verified here. It uses one fictional support ticket and no model invocation. The only intended downstream action is assigning a display-only review queue; it cannot refund, send mail, change an account, or call another system.

  1. Fix one question: Which existing review queue fits this ticket? Use exactly billing, shipping, and needs_review. Keep the ticket text, candidate descriptions, policy revision, chosen local model/card revision, PR/commit identity, and a deadline in one manifest.
  2. Before a decision, use ordinary code to verify that the ticket exists, the three queues exist, the policy revision is current, and the requester has no authority to create a fourth option. Those are facts, not decision-model work.
  3. Write three recorded fixture outcomes without calling the endpoint: a pass row where the typed answer is an allowed candidate and confidence meets the stated threshold; a reject row where it names an unknown candidate or has an invalid contract; and an unknown row where confidence is below the threshold. unknown means unavailable to this workflow, not a request to generate prose or select another backend.
  4. For a future separately authorized run, retain only non-sensitive state hash, selected option, full probability vector, confidence, model/card revision, input-token count, elapsed time, HTTP status, and policy result. Compare the answer with a human label on held-out synthetic cases before a display-only rollout.

The fixture fails closed for any 400 or 501, missing candidate, stale policy, timeout, malformed response, or low-confidence result. It does not invent a replacement route. A deterministic authorization check remains after the model signal: only code may permit the display assignment, and no model result may authorize a side effect.

The related Jev foundations chapter explains the general decision-design mental model. This update adds a dated local llama.cpp contract and the distinct code-versus-weight-license boundary. It does not compare the quality, cost, latency, calibration, or safety of Jev, OpenJev, Laya, Clef, or any hosted service. The forward-pass and typical-use descriptions in the announcement are publisher claims, not an independent benchmark.

The readable r/LocalLLaMA thread, posted October 2, is a community observation that links the same Hugging Face announcement. Its discussion raises classification versus autoregressive constraints, calibration, latency, and model-size questions. Reported hardware or latency anecdotes lack the exact commit, quantization, state, and test corpus needed for comparison. They are prompts for an isolated measurement plan, not evidence for a model ranking, performance figure, security property, or release fact.

MENTAL MODEL / REASONING ORDER

From an announcement to your own decision.

Primary sources

Compare the announcement with the conditions in the paper and official documentation.

Sources

Publication dates belong to the source; access dates record when it was checked. Community observations are separate from official statements.

01
New in llama.cpp: Decision Models ↗huggingface.coPublished: 2026-10-02 · Accessed: 2026-10-04
02
llama.cpp pull request #29818Published: 2026-10-01 · Accessed: 2026-10-04
03
llama.cpp commit a4cb4c6Published: 2026-10-02 · Accessed: 2026-10-04
04
llama.cpp server READMEPublished: Unknown · Accessed: 2026-10-04
05
OpenJev GGUF model card ↗huggingface.coPublished: Unknown · Accessed: 2026-10-04
06
llama.cpp MIT licensePublished: Unknown · Accessed: 2026-10-04
07
r/LocalLLaMA: New in llama.cpp Decision Models (community observation) ↗www.reddit.comPublished: 2026-10-02 · Accessed: 2026-10-04
Saved in this browser only.