A typed answer is not permission to act
On October 2, 2026, ggml-org published New in llama.cpp: Decision Models. It describes a local server endpoint, /v1/systemone, where a caller supplies a state and typed questions and receives probabilities over supplied options. The implementation trail is concrete: pull request #29818 was opened October 1, merged October 2, and produced commit a4cb4c6. The PR describes decision-model support as small, self-contained handling around embedding-style models and GGUF decision metadata.
This is a local runtime integration boundary, distinct from the earlier JSON-schema parser update. JSON Schema constrains generated syntax. /v1/systemone returns a choice, score, or yes/no-shaped judgement with model probabilities. Neither establishes that the judgement is true, current, calibrated for a business, or authorized to trigger a side effect.
- 1Observed facts and deterministic policy
- 2fixed candidate options
- 3one typed question
- 1Pinned local decision model
- 2probabilities and typed answer
- 3code checks threshold and authority
- 1Low confidence, unsupported model, invalid request, or stale policy
- 2unavailable result
- 3no mutation
Read the server contract before using a model
The current server README is an operating contract, not a dated release. It says questions are independent except for Clef's joint evaluation. Pin the model and revision, then inspect that model's option ceiling and truncation behavior before accepting a request; a shared endpoint is not an interchangeable classifier. A choice carries the highest-probability supplied option, distribution, and confidence; score is probability-weighted and noul is a true-probability. output_tokens is always zero, while stored temperature does not guarantee calibration on caller data. Invalid requests are 400; a non-decision model, or unsupported image input, is 501. Record these as unavailable contract states. Do not silently select text generation, another model, or a remote provider.
The llama.cpp repository license is MIT. It does not transfer to model weights. The inspected visible OpenJev GGUF card displays CC-BY-NC-4.0; its page creation time is not treated as a publication date. Check the specific card, revision, terms, and intended use before a download or deployment. This article did not download a model, accept a model license, start an endpoint, or make an API call.
An unexecuted N=1 offline fixture
This is an original worksheet, not a measured result or a server procedure verified here. It uses one fictional support ticket and no model invocation. The only intended downstream action is assigning a display-only review queue; it cannot refund, send mail, change an account, or call another system.
- Fix one question:
Which existing review queue fits this ticket?Use exactlybilling,shipping, andneeds_review. Keep the ticket text, candidate descriptions, policy revision, chosen local model/card revision, PR/commit identity, and a deadline in one manifest. - Before a decision, use ordinary code to verify that the ticket exists, the three queues exist, the policy revision is current, and the requester has no authority to create a fourth option. Those are facts, not decision-model work.
- Write three recorded fixture outcomes without calling the endpoint: a
passrow where the typed answer is an allowed candidate and confidence meets the stated threshold; arejectrow where it names an unknown candidate or has an invalid contract; and anunknownrow where confidence is below the threshold.unknownmeans unavailable to this workflow, not a request to generate prose or select another backend. - For a future separately authorized run, retain only non-sensitive state hash, selected option, full probability vector, confidence, model/card revision, input-token count, elapsed time, HTTP status, and policy result. Compare the answer with a human label on held-out synthetic cases before a display-only rollout.
The fixture fails closed for any 400 or 501, missing candidate, stale policy, timeout, malformed response, or low-confidence result. It does not invent a replacement route. A deterministic authorization check remains after the model signal: only code may permit the display assignment, and no model result may authorize a side effect.
The related Jev foundations chapter explains the general decision-design mental model. This update adds a dated local llama.cpp contract and the distinct code-versus-weight-license boundary. It does not compare the quality, cost, latency, calibration, or safety of Jev, OpenJev, Laya, Clef, or any hosted service. The forward-pass and typical-use descriptions in the announcement are publisher claims, not an independent benchmark.
The readable r/LocalLLaMA thread, posted October 2, is a community observation that links the same Hugging Face announcement. Its discussion raises classification versus autoregressive constraints, calibration, latency, and model-size questions. Reported hardware or latency anecdotes lack the exact commit, quantization, state, and test corpus needed for comparison. They are prompts for an isolated measurement plan, not evidence for a model ranking, performance figure, security property, or release fact.
MENTAL MODEL / REASONING ORDER
From an announcement to your own decision.
Compare the announcement with the conditions in the paper and official documentation.
Sources
Publication dates belong to the source; access dates record when it was checked. Community observations are separate from official statements.
01