A default is part of the production contract

On September 9, 2026, vLLM v0.29.0 made Model Runner V2 the default for all models. Its release notes also say that Model Runner V1 remains in use for some ROCm models and features that V2 does not yet support. That is a documented scope condition, not a reason to keep two serving designs alive. An operator adopting a current release needs to discover which runner actually starts for the intended model and configuration, record that result, and reject an unsupported target state rather than silently accepting an unexpected one.

This matters because a runner affects more than speed. It can change memory behavior, supported model paths, scheduling, and the shape of a failure. The v0.29.0 notes also introduced queue-admission controls and listed removals and deprecated entrypoints. A deployment that starts after an upgrade can still be wrong if its old environment variable is ignored, its chosen backend differs from the recorded target, or its queue reaches a resource limit before a caller receives a controlled rejection.

  1. 1Release tag and image digest
  2. 2intended model and runner
  3. 3explicit queue limits
  1. 1Route inventory
  2. 2allow or deny decision per route
  3. 3isolated fixture
  1. 1Observed runner and responses
  2. 2adopt the pinned release or reject the target state
Consider the sequence and each role.

The source repository is licensed under Apache-2.0. That establishes the repository code's stated license; it does not establish rights for a model, a checkpoint, a container dependency, or an input dataset. Record those as separate artifacts before a serving decision.

Version 0.30 changes the exposed surface too

The v0.30.0 release, published September 22, changes more than an internal implementation. It says that scale-out endpoints are not registered by plain vllm serve unless --enable-scale-out is passed. It also removes the VLLM_ENABLE_SCALE_OUT_ENDPOINTS environment-variable route, lists other removed or deprecated configuration paths, and deprecates the module-form gRPC entrypoint in favor of vllm serve --grpc.

Those statements suggest a precise review question: which routes exist in the target process, and which caller is permitted to reach each one? Do not infer the answer from a process-level key. An earlier v0.28.0 release, dated August 26, says documentation was updated to warn that --api-key does not gate all endpoints. This historical warning is outside the current research window; it is not a new v0.30.0 change. It motivates a separate route-policy review: inspect network placement, gateway authorization, and the application’s exposure choices together. If an intended route has no verified authorization owner, leave it unavailable.

The repository's security policy provides a private reporting path and describes issue triage. It is useful when handling a suspected vulnerability, but it does not certify a deployment, enumerate every reachable endpoint, or replace a local authorization test.

An offline adoption fixture to run before a change

This is an unexecuted exercise, not a verified vLLM procedure. Use an isolated, non-sensitive environment and the intended pinned release artifact. First, write a small manifest containing the release tag, resolved image digest or package hash, model revision, hardware/backend, intended runner, network listeners, and each enabled route. Do not use a mutable image tag as the identity of a tested artifact.

Start only the intended configuration. Capture its startup output and identify the runner that was actually selected. Send one representative permitted request and save only non-sensitive request metadata and the resulting status. Then exercise the queue boundary with a deliberately small, documented limit. The acceptance condition is not a throughput number: it is a bounded, observable response when the limit is reached.

Next, make a route table. For every discovered route, test one caller that should be permitted and one that should be denied. A denied request should fail closed: no model execution, no partial administrative action, and no ambiguous success response. Include the scale-out routes in the table even when they are absent by default; absence is a meaningful observed state. A request that unexpectedly succeeds is a failed fixture, not an invitation to add an unreviewed exception.

Finally, compare the observations with the manifest. Do not retain an old runner as a compatibility path, rely on automatic rollback, or call a documented exception a successful migration. Pin the intended current release, correct the configuration or authorization boundary, and rerun the isolated fixture until the target state is explicit. No benchmark or security result is asserted here; the release notes describe vendor-maintained software changes, while this exercise defines evidence an operator would still need.

MENTAL MODEL / REASONING ORDER

From an announcement to your own decision.

Primary sources

Compare the announcement with the conditions in the paper and official documentation.

Sources

Publication dates belong to the source; access dates record when it was checked. Community observations are separate from official statements.

01
vLLM v0.28.0 release: endpoint authentication warningPublished: 2026-08-26 · Accessed: 2026-10-04
02
vLLM v0.29.0 releasePublished: 2026-09-09 · Accessed: 2026-10-04
03
vLLM v0.30.0 releasePublished: 2026-09-22 · Accessed: 2026-10-04
04
vLLM Apache-2.0 licensePublished: Unknown · Accessed: 2026-10-04
05
vLLM security policyPublished: Unknown · Accessed: 2026-10-04
Saved in this browser only.