A tool-call close must also end its in-memory identity
llama.cpp b11393, published on October 4, 2026, is a prerelease. Its linked pull request #29942 was created at 10:12 UTC and merged at 15:59 UTC. The change concerns the chat-PEG parser, not JSON Schema output or tool authorization in general.
The pull request describes a state-lifetime error: TOOL_CLOSE reset the pending tool call while current_tool still referenced that destroyed object. A later TOOL_ID could then act through the stale reference. The merged commit dbe4c3e clears current_tool when the pending call is reset. The author reports use-after-free and double-free with identifiers in different allocator size classes, and says a regression case was added. The commit's changed-file list inspected here contains only the parser repair; it does not establish that the reported test was merged. These are author reports, not an independently reproduced security assessment, a CVE, or a guarantee across parsers.
- 1TOOL_OPEN
- 2pending tool call and current_tool
- 3TOOL_CLOSE
- 4copy completed call
- 1TOOL_CLOSE
- 2reset pending call
- 3clear current_tool
- 4later TOOL_ID has no live target
- 1A parser/template that cannot express this order
- 2outside this exact evidence
The current server README describes function calling and tool use as server features, but it does not turn a parser repair into permission control. A valid tool-call shape is still separate from whether an application allows the named operation, what credentials it holds, and whether its listener is reachable. The repository is MIT-licensed code; model weights, tools, prompts, and deployment data have their own terms. Its security policy also advises against using server functionality on an untrusted network when the model cannot run in a secure, isolated environment.
A specific path from 2024 to the current mapper
PR #9639 was opened on September 25, 2024 and merged on January 30, 2025, bringing generic and model-native tool-call parsing into the server. PR #18675, merged on March 6, 2026, refactored the architecture around PEG parsers and template analysis. These dated changes explain why a template and a parser revision belong in a test record: the wire request alone does not identify the code path. They do not show that every model uses the same grammar or that b11393 fixes an older downstream UI.
Pin the grammar path, then test the prohibited transition
Treat b11393 as a candidate to evaluate, not a reason to roll a local endpoint forward blindly. Record the release tag, resolved commit, platform artifact checksum, model revision, chat template, parser selection, request bytes, and listener scope. A test that only sees one successful tool call cannot observe the transition this patch addresses.
This is an original, unexecuted N=1 fixture for one local consumer:
- In an isolated, loopback-only environment, use one pinned b11393 build and one parser/template combination. Do not provide tools with filesystem, browser, network, or credential access.
- Feed a minimal parser fixture whose token order is
TOOL_OPEN, an identifier A,TOOL_CLOSE, then identifier B. Keep it as a parser test artifact, not a model-generated command. Capture exit status, parser diagnostic, build identity, and elapsed time with a fixed timeout. - The pass condition is narrow: the process remains healthy and the post-close identifier is not attached to the completed call. A crash, sanitizer report, unexpected attachment, timeout, or different parser identity is a rejected result. Preserve the artifact and inspect that one boundary before changing inputs.
This fixture describes the state contract. It does not specify allocator size classes and therefore does not reproduce the author's memory-fault setup. Before implementing it, inspect the pinned common/chat-peg-parser.cpp mapper and tests/test-chat-peg-parser.cpp harness, then encode the events using that revision's actual parser rules. TOOL_OPEN and TOOL_ID here are internal event labels, not request fields or literal model tokens. Keep a trace row for each event: event index, pending-call presence, whether a live current-tool reference exists, and completed-call identifier. The final row must preserve A on the completed call while B has no closed-call target. If the selected grammar cannot produce that order, record an unsupported fixture instead of claiming a pass.
Do not replace failure with a second runtime, an unconstrained parser, automatic retry, or tool execution. The fixture does not measure exploitability, remote reachability, model behavior, throughput, memory safety across templates, or security of a multi-tenant service. No download, build, parser run, checksum check, sanitizer run, or benchmark was performed for this note; any such work consumes local compute, storage, and operator time.
Why this is separate from schema-constrained output
The earlier schema-boundary note addresses the Ling 3.0 response-format grammar path. This b11393 evidence instead tracks the lifetime of a chat-PEG tool-call object after its close marker. Both argue for pinning a parser and testing one consumer contract, but neither lets a result transfer to another parser family. For an application that must invoke a real operation, keep a separate authorization decision after parsing; parser acceptance is only syntax and state handling.
MENTAL MODEL / REASONING ORDER
From an announcement to your own decision.
Compare the announcement with the conditions in the paper and official documentation.
Sources
Publication dates belong to the source; access dates record when it was checked. Community observations are separate from official statements.