A speculative decoder does not automatically have a fast batch path
The b11404 llama.cpp prerelease was published on October 5, 2026. Its merged change, pull request #29869, adds a Metal matrix-multiplication route for a narrow situation: a small number of activation rows can arise when a speculative decoder verifies draft tokens, or when batched decoding has only a few active sequences. That is a shape observation, not a promise that enabling speculative decoding makes every Apple device or model faster.
The author describes the previous route as a matrix-vector-style path whose work grows as those activation rows increase. The new route uses 8-by-8 simdgroup matrix tiles and shares each decoded weight over several rows. That can remove repeated dequantization work when there are enough rows to share it across, while avoiding a large matrix-matrix dispatch for a short verification batch.
- 1Draft tokens or a few active sequences
- 2activation rows in src1
- 1Eligible device + eligible layout + row threshold
- 2few-row Metal MMA route
- 1Any failed predicate
- 2existing dispatch path
- 1Pinned fixture
- 2compare correctness and latency within one recorded environment
This diagram is a conceptual reading guide, not a profile of a particular model. The PR's reported measurements are the author's measurements on one M3 Ultra, macOS version, model, quantization, and baseline. They are useful for explaining why thresholds exist, but they are neither an independent benchmark nor a result for another GPU, model, driver, build, prompt, or draft model.
Verification creates work before it removes decoding steps
Medusa v1, published January 19, 2024, illustrates a related scheduling idea: extra decoding heads propose continuations, and tree attention processes candidates before accepting a prefix. Its section 3.1.2 explains the tradeoff between more candidates and more verification computation. This is history for the learner's question about verification work, not evidence that b11404 implements Medusa or that their performance results transfer. The October 2026 patch concerns a backend operation; a candidate-generating algorithm still has its own acceptance rate and overhead.
The row count is only one gate
The implementation makes the decision from several properties. The activation tensor must be F32, untransposed, and have a compatible contiguous stride; its second dimension must be between a type-dependent minimum and 16 rows. The supported weight formats are a listed set of floating and quantized ggml types, and the K dimension must meet that format's step and layout requirements. Batch-shape values also have bounds because they are passed to Metal function constants.
The device condition matters equally. The patch selects this route on Apple GPU family 7 or later when the tensor API is not in use. The PR describes thresholds that it selected from its M3 Ultra measurements: six rows for F32, three for F16/Q4_K/Q5_0/Q5_1, and two for other supported types. Those numbers are selection policy in this revision, not a portable hardware law. A later Apple GPU can choose a different default capability path; a device, compiled binary, or tensor layout that does not meet the predicates continues through another implemented path. It is not a fallback added by this lesson.
The patch also permits a fused matrix-multiply-plus-same-shape-add operation under its own graph conditions. Do not infer that every residual add is fused. The added tests cover selected row counts, weight formats, layouts, batch relations, and graph ordering. They do not prove every model graph, every Metal driver, or end-to-end serving correctness.
Design one bounded comparison before changing a serving decision
Use this original, unexecuted N=1 exercise to decide whether the new branch is relevant to one local setup. It is deliberately not an install guide, download instruction, or performance result.
- Write a manifest before running anything: release
b11404, merge commita3a1c4747fdc0dcad40b3946108b89375d9a7d0e, resolved binary digest, macOS version, Apple GPU family/capability observation, build flags, model revision, weight format, draft configuration, prompt shape, number of active sequences, and timeout. Keep inputs synthetic and non-sensitive. - Create one inside-range and one outside-range case for the revision's 2–16 row window, using an explicitly recorded draft/batch setting to vary the rows. Within each case, compare the previous observed release b11401 with b11404 while holding the model, prompt, row setting, output limit, and concurrency fixed. Resolve and record both build digests before comparison. Record the expected routing state as unknown until a suitable local observation is collected; do not derive it from a release tag.
- In a separately authorized local run, record correctness criteria before timing: completion status, generated-token comparison policy, device/runtime diagnostics, and bounded elapsed-time samples. The author report is not a tolerance specification. Select any floating-point comparison tolerance before looking at the result and retain it in the manifest.
- Stop and keep the diagnostic boundary if the artifact identity differs, the device/capability cannot be established, a layout predicate is not observable, an output comparison fails, a process exits abnormally, or the timeout expires. Do not replace the model, driver, backend, or row shape and call that a result for the rejected case.
This exercise can consume local machine time, energy, model storage, and operator attention. The project README lists Metal as an Apple Silicon backend, but it does not establish that a particular binary, weights license, or draft configuration is suitable for your use. The repository is MIT-licensed; selected model weights and downloaded artifacts have their own terms. Its security policy advises isolation for untrusted models and cautions against exposing server functionality on untrusted networks. None of those documents turns this prerelease into a security certification.
Keep the decision narrower than “speculation is faster”
b11404 establishes a prerelease implementation and a maintainer test/measurement report. It does not establish an end-to-end speedup for this archive, a general Metal performance guarantee, compatibility with every Apple GPU, or the quality of a speculative draft model. The practical question is: does this pinned workload create an eligible short-row operation on this recorded Metal configuration, and does it preserve the correctness rule chosen before measurement?
That question is separate from the x86 K-tail shape boundary in the related x86 article and from router process framing in the router article. The 2024 Medusa source explains why verification work is a meaningful cost; only the pinned 2026 PR and implementation establish this Metal dispatch predicate. The research window was September 5, 2026 15:35 JST to October 5, 2026 15:35 JST.
MENTAL MODEL / REASONING ORDER
From an announcement to your own decision.
Compare the announcement with the conditions in the paper and official documentation.
Sources
Publication dates belong to the source; access dates record when it was checked. Community observations are separate from official statements.