One allocation can depend on the wrong shape dimension

The October 4, 2026 b11390 llama.cpp prerelease names a CUDA MMQ memory-fault fix for the condition n_expert >> n_ubatch. Its linked pull request, #29941, was created at 09:15 UTC, merged at 12:10 UTC, and produced commit dd266785. The author describes an illegal CUDA memory access in a specific large-tensor test configuration. This is a maintainer report attached to a prerelease, not a CVE, a security assessment, or evidence that every mixture-of-experts model has the same fault.

The useful mechanism is narrow. MMQ needs temporary q8_1 storage. For dense layouts, the maximum tile width can be bounded by one tensor dimension. The pull request says the MoE layout instead requires the dimension that reflects the expert count. Before the change, the allocation added padding computed from ne11; commit dd266785 uses ne12. When the number of experts is much larger than the physical micro-batch, the old bound can be too small and a CUDA kernel can access memory outside the temporary allocation.

  1. 1Pinned MoE tensor shape: n_expert, n_ubatch, quantization, backend
  2. 2temporary q8_1 allocation
  1. 1Correct layout dimension
  2. 2enough tile-width padding
  3. 3MMQ kernel observation
  1. 1CUDA illegal-memory-access, timeout, or mismatched manifest
  2. 2stop the fixture
  3. 3preserve diagnostic evidence
Consider the sequence and each role.

This does not make n_expert a universal safety switch. The exact tensor layout, effective expert count, quantization type, GPU compute capability, CUDA runtime, compiled artifact, model, and batch shape determine whether this particular path is even relevant. The same pull request says the conditions were highly specific and deliberately did not add the large reproduction tensors to test-backend-ops.cpp. A release tag alone cannot fill that coverage gap.

Design one isolated observation before changing a serving path

Use the following as an original, unexecuted N=1 fixture design. It is not a command sequence, a tested reproduction, or a recommendation to install software or download a model.

  1. Write one manifest for the intended failure boundary: b11390, commit dd266785c2595775001c1c714bd9d92b3ef34cde, resolved artifact digest, CUDA driver/runtime, GPU model and compute capability, exact MoE model revision, quantization, n_expert, n_ubatch, other tensor dimensions, request shape, and a short timeout. Keep the fixture on an isolated host and use synthetic, non-sensitive input only.
  2. Give a dedicated operator permission to start and stop that one local process. Do not give it production credentials, network listeners, tool execution, shared model caches, or access to customer data. The repository's MIT license describes its code; it does not establish rights to any selected model weights, dataset, or driver component.
  3. In a separately authorized future run, capture the artifact identity, manifest, exit status, CUDA error text, elapsed time, and a bounded diagnostic log. A normal exit only says this one recorded case did not show the observed failure. It does not establish throughput, memory efficiency, correctness, or coverage of another shape.
  4. Stop immediately on an illegal-memory-access report, watchdog timeout, device reset, artifact mismatch, or loss of the manifest. Mark the configuration unavailable for this consumer and diagnose one changed condition at a time. Do not retry it through another backend, substitute another model, or silently lower the workload until the original boundary disappears.

The cost is local GPU time, electricity, operator time, temporary storage, and potential interruption of the isolated machine. Set a duration and log-size limit before the run. No paid API is required by this design, but no cost estimate or execution result was produced here.

Keep the claim smaller than the fix name

The b11390 release and its merge establish that the project shipped a prerelease containing the one-line dimension change. They do not establish exploitability, CVE assignment, a general CUDA security property, compatibility with every MoE model, or an improvement in latency. The current CUDA source and repository security policy are undated operating material; neither is recent-release evidence. They were read only to identify the code and project boundary.

There is no necessary 2024–2025 history source for this note. A broad CUDA or MoE timeline would not explain this allocation decision. The relevant lineage begins with the October 4 pull request, merge, and prerelease. For a different boundary, see the related parser-specific structured-output fixture and local decision-endpoint fixture. Neither evaluates GPU memory behavior.

MENTAL MODEL / REASONING ORDER

From an announcement to your own decision.

Primary sources

Compare the announcement with the conditions in the paper and official documentation.

Sources

Publication dates belong to the source; access dates record when it was checked. Community observations are separate from official statements.

01
llama.cpp b11390 prereleasePublished: 2026-10-04 · Accessed: 2026-10-04
02
llama.cpp pull request #29941Published: 2026-10-04 · Accessed: 2026-10-04
03
llama.cpp commit dd266785Published: 2026-10-04 · Accessed: 2026-10-04
04
llama.cpp CUDA backend sourcePublished: Unknown · Accessed: 2026-10-04
05
llama.cpp MIT licensePublished: Unknown · Accessed: 2026-10-04
06
llama.cpp security policyPublished: Unknown · Accessed: 2026-10-04
Saved in this browser only.