The affected contract is smaller than “JSON output works”

llama.cpp b11377, published on October 3, 2026, is a prerelease rather than a broad compatibility promise. Its linked change, pull request #29813, says that the Ling 3.0 parser previously built a grammar for tool calls but did not handle inputs.json_schema; the release describes affected response_format requests as unconstrained. The merged commit 9bf55f4 adds a response-format grammar path ahead of tools, requires a closing thinking block before JSON when thinking is enabled, and disallows trailing prose after JSON.

That is a parser-specific reliability change. It does not prove that every llama.cpp model, template, API client, JSON Schema feature, tool call, or endpoint now produces valid business data. The current server README documents both plain JSON and schema-constrained response_format, but also limits its compatibility language: it says the chat-completions endpoint suffices for many applications in its experience, and that models need a supported chat template for optimal use. “One request looked like JSON” remains weaker evidence than parsing and independently validating the exact schema needed by one consumer.

  1. 1Pinned b11377 plus Ling 3.0 template
  2. 2response_format with one named schema
  3. 3raw response
  1. 1Independent JSON parser
  2. 2schema validation
  3. 3accept the one consumer result
  1. 1Invalid JSON, extra prose, timeout, or schema mismatch
  2. 2reject the target state
  3. 3diagnose one boundary
Consider the sequence and each role.

The relevant dated history is deliberately narrow. Structured output has been an active interface problem since at least 2024, but this article does not reuse a general industry timeline as evidence for a local runtime. For this specific path, the observable lineage is the October 1, 2026 pull request, its October 3 merge and the October 3 b11377 prerelease. That is enough to identify what changed; it is not evidence that an older or different parser had the same behavior.

Pin the parser, model, and template together

The b11377 release points at build artifacts and an attestation, but a release tag is still not the complete identity of an observed run. Before deciding to use it, record the b11377 tag, resolved artifact checksum, commit 9bf55f4a3677af697d914d959eaa70f93cfdc494, model revision, exact Ling 3.0 chat-template configuration, endpoint, request payload, response-format schema, backend, and platform. A model change can alter its template behavior; a template change can alter where reasoning text appears; a client can parse a response differently from the server. Combining all of those facts into an anonymous “structured output version” loses the failure boundary.

The repository declares an MIT license, which identifies the repository code's stated terms. It does not grant rights to a model, weight file, dataset, client library, or prompt. The security policy also cautions that its server functionality should not be used on an untrusted network and excludes some experimental or untrusted-environment surfaces from covered security topics. It is a project policy, not a security assessment of b11377 or of a local service.

Keep this fixture local and non-sensitive. Do not enable agent or tool execution; do not give it a filesystem, browser, credentials, or an unreviewed remote MCP command. The server README states that tool or agent modes can expose file read/write through the API and changes the CORS default. A JSON constraint governs emitted syntax; it does not authorize operations, prevent prompt injection, or isolate a model process.

An unexecuted, bounded fixture

This is an original exercise, not a tested b11377 procedure or a claimed conformance result. It has one concrete consumer and intentionally introduces no general schema wrapper.

  1. In an isolated environment, obtain a verified b11377 artifact and pin its checksum. Start only the intended Ling 3.0 configuration on a loopback listener. Record a short client timeout, for example 10 seconds, and a fixed request-size limit before sending any request. A timeout is a failed observation, not permission to retry against a different parser.
  2. Define one small schema that the consumer actually needs, such as an object with a required string decision limited to allow or deny, and a required integer confidence from 0 through 100. Keep a valid fixture request and an intentionally invalid or contradictory request separate. Neither is a production decision.
  3. Submit each request with the documented schema-constrained response_format. Save non-sensitive raw bytes, HTTP status, elapsed time, exact server/template identity, and validator output. First parse as JSON, then validate against the same schema outside the model client. Do not accept text merely because a regular expression finds braces.
  4. Fail the fixture if the client times out, the server returns a non-success status, parsing fails, extra prose is present where the consumer forbids it, a required field is absent, an enum is outside the schema, or the recorded parser/template differs from the manifest. Preserve the evidence and classify the failure as transport, server, template, generation, parser, or validator before changing one condition.

The acceptance result is only “this pinned configuration produced schema-valid output for these two recorded inputs.” It is not accuracy, safety, tool authorization, throughput, or cross-model compatibility. A failed fixture should leave the feature unavailable to that consumer; it should not silently switch to unconstrained JSON, a different chat template, a second runtime, or an automatic repair path.

This exercise may consume local compute time, electricity, storage, and staff time. It calls for no paid API and no external model upload, but that does not make its operational cost zero. Stop within the stated timeout and resource limit, retain only non-sensitive artifacts, and remove the isolated test environment according to its own data-handling policy. No benchmark, schema-conformance run, checksum verification, or security test was performed while writing this note.

MENTAL MODEL / REASONING ORDER

From an announcement to your own decision.

Primary sources

Compare the announcement with the conditions in the paper and official documentation.

Sources

Publication dates belong to the source; access dates record when it was checked. Community observations are separate from official statements.

01
llama.cpp b11377 releasePublished: 2026-10-03 · Accessed: 2026-10-04
02
llama.cpp pull request #29813Published: 2026-10-01 · Accessed: 2026-10-04
03
llama.cpp commit 9bf55f4Published: 2026-10-03 · Accessed: 2026-10-04
04
llama.cpp server READMEPublished: Unknown · Accessed: 2026-10-04
05
llama.cpp MIT licensePublished: Unknown · Accessed: 2026-10-04
06
llama.cpp security policyPublished: Unknown · Accessed: 2026-10-04
Saved in this browser only.