Keep authority with the deployment

vLLM v0.31.0, published October 5, 2026, lists PR #58830, created and merged September 26. The PR adds a default-false --trust-request-mm-kwargs gate. In its inspected online-rendering and pooling/scoring paths, non-empty request mm_processor_kwargs or media_io_kwargs raise VLLMValidationError unless that flag is enabled. The author says these values can change media loading and preprocessing resource use; server-level settings and offline LLM use remain unchanged.

  1. 1Operator: deployment media policy
  2. 2server configuration
  3. 3bounded default
  1. 1Client: text + media
  2. 2covered request validator
  3. 3empty kwargs proceed to normal validation
  1. 1Client: non-empty request kwargs
  2. 2author-tested validation rejection
  1. 1Trusted internal client + explicit server flag
  2. 2override eligible for a separate test
Consider the sequence and each role.

This is an authority map, not an endpoint survey. The patch's tests cover default rejection, opt-in acceptance, CLI wiring, and one flash late-interaction path; this article did not run them. It establishes neither a safe media size nor a CVE, severity, model/backend scope, or complete API boundary. The repository declares Apache-2.0 for code; weights, media, containers, and datasets need separate review. Its security policy describes reporting, not deployment certification.

The previous v0.29–v0.30 lesson concerns runner and route exposure. This one asks who may select preprocessing work after a permitted request arrives. Do not enable the flag merely because a feature wants a different frame count or processor setting: place an approved value in the deployment manifest first. An internal client is a defined identity and constraint set, not a name, subnet, or successful request.

An offline rejection matrix before enabling anything

This is an original, unexecuted N=1 fixture. It does not install vLLM, fetch a model, start a server, or send an API request. Use it to write the acceptance criteria before an isolated test environment is allowed to consume accelerator time.

Case Server flag Request fields Expected classification
A absent both omitted or empty candidate for normal request validation
B absent non-empty mm_processor_kwargs request validation failure; record field name
C absent non-empty media_io_kwargs request validation failure; record field name
D present one approved non-empty override only eligible for a trusted-client test

For a real isolated run, pin v0.31.0 and record the resolved artifact digest, model revision, hardware/backend, listeners, authenticated caller, server-level media settings, the single requested override, and a timeout plus memory/CPU/GPU budget. Run B and C first. Treat an accepted override without the flag, preprocessing beginning after a rejection, an unbounded queue, or missing field-level diagnostics as a failed fixture. Stop and inspect the release, configuration, and reverse-proxy request handling; do not add a second backend or a permissive fallback.

Case D is not an automatic pass. It requires a defined internal identity, a narrow approved value, and a cost limit. The permitted request may still decode a large file, wait on media I/O, allocate processor state, or fail for model-specific reasons. Log only non-sensitive request metadata and resource counters; media and prompts can be confidential. A timeout or budget breach is evidence to tighten the deployment policy, not evidence that the trust flag should remain enabled.

History and decision boundary

The September 5, 2024 vLLM performance post describes API-server CPU-work isolation from the inference engine. That reported performance design neither transfers a speed result nor establishes the 2026 client-authority gate. The dated request-authority evidence remains the September 26 PR and October 5 release.

Adopt the gate as a default-deny boundary when an API is reachable by callers whose preprocessing choices are not under the operator's control. Consider the explicit opt-in only when the caller, allowed override, resource budget, and diagnostic path can all be tested together. This article makes no benchmark, availability, CVE, endpoint-completeness, or independent security claim; its fixture remains unexecuted.

MENTAL MODEL / REASONING ORDER

From an announcement to your own decision.

Primary sources

Compare the announcement with the conditions in the paper and official documentation.

Sources

Publication dates belong to the source; access dates record when it was checked. Community observations are separate from official statements.

01
Official documentationvLLM v0.31.0 releasePublished: 2026-10-05 · Accessed: 2026-10-05
02
Official documentationvLLM pull request #58830: Gate per-request multimodal processor kwargsPublished: 2026-09-26 · Accessed: 2026-10-05
03
Official blogvLLM v0.6.0 performance update ↗vllm.aiPublished: 2024-09-05 · Accessed: 2026-10-05
04
Official documentationvLLM Apache-2.0 licensePublished: Unknown · Accessed: 2026-10-05
05
Official documentationvLLM security policyPublished: Unknown · Accessed: 2026-10-05
Saved in this browser only.