Vulnerability GHSA-ph72-cqr5-qpp7
Summary
vLLM: Scale-out disaggregated multimodal transport trusts caller-supplied features
Details
Affected
- Ecosystem / package: pip /
vllm - Affected versions: vLLM ≤ 0.25.1 (confirmed on 0.25.1, commit
752a3a504485). The lower bound predates 0.25.1; maintainers can confirm how far back the scale-out transport path reaches.
Summary
vLLM's disaggregated scale-out transport splits a multimodal request into a trusted render step (POST /v1/chat/completions/render) and a separate generate step (POST /inference/v1/generate). The generate route decodes a caller-supplied features object — serialized encoder tensors (kwargs_data), multimodal hashes (mm_hashes), placeholder ranges (mm_placeholders), and the internal field-processor selection — and forwards it into the engine as if it had come from the trusted renderer, with no rebinding to (or validation against) the active model's renderer contract. Because the two routes are ordinary auth-guarded HTTP endpoints (the /inference prefix is registered by default on generate-capable servers), any authenticated caller can submit an otherwise-valid render body with a single forged field.
Depending on which field is forged, this produces:
- an engine-fatal crash of the shared EngineCore process (denial of service), reproduced as a CUDA illegal-memory-access, a post-admission rank-mismatch
ValueError, and a hardassert— three independent forged fields (sites 1, 2, 3); - silent cross-request encoder-cache poisoning / disclosure of shared encoder state when the cache-key hash is not bound to the payload (site 4);
- transport-level integrity loss when sparse placeholder masks are dropped during render-to-generate replay (site 5).
All five share one root cause and one fix shape: the reconstructed multimodal state on the scale-out path is trusted without being rebound to, and validated against, the active model's renderer output before it reaches the engine.
These sites are distinct from prior multimodal hardening. Site 1 survives GHSA-wv77-2vpf-vmmg (that fix validates full tensor shape in MultiModalDataParser/get_input_embeddings on the prompt-embeds path), because our request forges image_grid_thw metadata with the pixel bytes intact and reaches the Qwen2 vision RoPE/cu_seqlens and image_embeds.split sink, which the shape-check fix does not rebind. Site 4 is distinct from GHSA-c65p-x677-fgj6 (which folds metadata into MultiModalHasher.serialize_item to stop hash collisions), because the scale-out generate path trusts a caller-supplied mm_hash as the cache key with no origin binding, so that fix does not stop a caller from submitting a victim's hash or a kwargs_data=None cache read.
Affected code
Links pinned to the confirmed commit 752a3a504485 (v0.25.1).
Shared entry point and control surface for all five sites:
POST /inference/v1/generateroute:vllm/entrypoints/scale_out/token_in_token_out/api_router.py#L46-L75.- Public
featuresschema (kwargs_data,mm_hashes,mm_placeholders):vllm/entrypoints/scale_out/token_in_token_out/protocol.py#L42-L63. ServingTokens.serve_tokens()copies the decoded geometry and hashes into engine structures without rebinding to the renderer schema:vllm/entrypoints/scale_out/token_in_token_out/serving.py#L145-L172.- The scale-out routers are registered by default on generate-capable servers and
/inferenceis treated as an ordinary auth-guarded prefix:vllm/entrypoints/openai/api_server.py#L217-L219.
Site 1 — forged Qwen grid geometry (engine-fatal DoS). The decoded image_grid_thw is never rebound to the rendered pixel-tensor element count.
- Reconstruction with no model-specific geometry invariant:
serving.py#L147-L172. Qwen2VisionTransformer.prepare_encoder_metadata()derives RoPE tables,cu_seqlens, and the FlashAttentionmax_seqlenfrom the caller-supplied grid:qwen2_vl.py#L658-L712, consumed inforward()(#L713-L751)._process_image_input()computesimage_embeds.split(sizes)from the same untrusted metadata:qwen2_vl.py#L1348-L1369; M-RoPE positions atqwen2_vl.py#L1223.
The only shape check is assert grid_thw.ndim == 2; the split sizes and the vision-encoder call are then derived directly from the caller-supplied grid, with no cross-check against the pixel-tensor row count:
# vllm/model_executor/models/qwen2_vl.py Lines 1348-1369
def _process_image_input(
self, image_input: Qwen2VLImageInputs
) -> tuple[torch.Tensor, ...]:
grid_thw = image_input["image_grid_thw"]
assert grid_thw.ndim == 2
if image_input["type"] == "image_embeds":
image_embeds = image_input["image_embeds"]
else:
pixel_values = image_input["pixel_values"]
if self.use_data_parallel:
return run_dp_sharded_mrope_vision_model(
self.visual, pixel_values, grid_thw.tolist(), rope_type="rope_3d"
)
else:
image_embeds = self.visual(pixel_values, grid_thw=grid_thw)
# Split concatenated embeddings for each image item.
merge_size = self.visual.spatial_merge_size
sizes = (grid_thw.prod(-1) // merge_size // merge_size).tolist()
return image_embeds.split(sizes)
Site 2 — wire-selected field-processor type confusion (engine-fatal DoS). MsgpackDecoder._decode_mm_field_elem() trusts a wire-selected field-factory name and constructs the internal field processor directly from caller data.
vllm/v1/serial_utils.py#L440-L454readsfactory_meth_name, factory_kw = obj["field"]and callsgetattr(MultiModalFieldConfig, factory_meth_name).- Qwen2-VL field/schema contract:
qwen2_vl.py#L763-L791andQwen2VLImagePixelInputs(#L119-L144); parse/validate atqwen2_vl.py#L1300. - Rank mismatch raised post-admission:
vllm/utils/tensor_schema.py#L155-L171, reached fromTensorSchema.__init__→validate()(#L63).
Site 3 — non-positive placeholder length (engine-fatal DoS via reachable assert). PlaceholderRangeInfo{offset,length} is accepted as unconstrained integers and copied verbatim into the engine's PlaceholderRange.
- Schema:
protocol.py#L28-L35. - Copied into
PlaceholderRange:serving.py#L147-L153. - The only guard is an upper bound on embed count (
num_embeds > mm_encoder_cache_size), with no non-positive check:vllm/v1/engine/input_processor.py#L459-L464; raw length returned byvllm/multimodal/inputs.py#L152-L154; window selection assumes non-empty ranges atvllm/multimodal/utils.py#L114-L134. - Sink:
assert start_idx < end_idxatvllm/v1/worker/gpu/mm/encoder_runner.py#L114(duplicated atvllm/v1/worker/gpu_model_runner.py#L3192); turned into a fatal shutdown by EngineCore's uncaught-exception path atvllm/v1/engine/core.py#L1229-L1233.
A length of 0 makes num_encoder_tokens == 0, so end_idx collapses to 0 and the bare assert fires inside the engine worker:
# vllm/v1/worker/gpu/mm/encoder_runner.py Lines 108-114
pos_info = mm_feature.mm_position
start_pos = pos_info.offset
num_encoder_tokens = pos_info.length
start_idx = max(cur_query_start - start_pos, 0)
end_idx = min(cur_query_end - start_pos, num_encoder_tokens)
assert start_idx < end_idx
Site 4 — cache hash not bound to payload (integrity / disclosure). features.mm_hashes (the cache key) and kwargs_data (the tensor) are independent fields with no origin or integrity binding.
- Schema exposing the caller-controlled hash and the
None= resolve-from-cache semantics:protocol.py#L42-L63. - Handler forwards the caller's hashes unchanged into
mm_input(...):serving.py#L164-L170. - Input processing copies the caller hash into the feature identifier with only a string-type check:
vllm/v1/engine/input_processor.py#L165-L181. - Sink — the receiver cache returns the cached tensor solely by that key (
cache_key = feature.mm_hash or feature.identifier):vllm/multimodal/cache.py#L602-L607.
The cache key is the caller-supplied hash with no verification against the tensor bytes, so a forged mm_hash both stores under and reads back another request's slot:
# vllm/multimodal/cache.py Lines 601-607
for feature in mm_features:
cache_key = feature.mm_hash or feature.identifier
self.touch_receiver_cache_item(cache_key, feature.data)
for feature in mm_features:
cache_key = feature.mm_hash or feature.identifier
feature.data = self.get_and_update_item(feature.data, cache_key)
return mm_features
Site 5 — dropped sparse placeholder mask (transport integrity loss). The render path serializes placeholders as only offset/length, so models relying on sparse is_embed masks lose the mask during render-to-generate replay.
ServingRender._extract_mm_features()builds eachPlaceholderRangeInfo(offset=p.offset, length=p.length), discardingis_embed:vllm/entrypoints/scale_out/render/serving.py#L212-L229.- Transport schema carries no field for the mask:
protocol.py#L28. ServingTokens.serve_tokens()reconstructs a densePlaceholderRangeregardless of the original:serving.py#L148-L152.
Impact
A single authenticated request to a scale-out multimodal deployment can:
- Crash the shared EngineCore process (sites 1, 2, 3), taking the served model down for every tenant (
/health→ 503). Availability-only; no code execution or data disclosure demonstrated for these sites. - Silently poison or read back another request's shared encoder-cache state (site 4) — an integrity/disclosure primitive. Attack complexity is High because the attacker must know or induce the victim's content hash; there is no availability impact for this site.
- Corrupt backend-visible placeholder semantics across the render-to-generate boundary (site 5) for models that depend on sparse
is_embedmasks.
The forged multimodal payload is small; only the trust in its self-declared geometry/identity is the defect.
Suggested Fix
On the scale-out path, do not trust caller-supplied multimodal state as renderer-produced. After decoding features, rebind and validate the reconstructed MultiModalKwargsItem against the active model's renderer contract at the HTTP boundary:
- Reject any request whose decoded grid geometry is inconsistent with the rendered pixel-tensor element count and declared placeholder span, before it reaches
prepare_encoder_metadata()(site 1). - Rebind each field's processor type to the schema the active model's renderer declares (or reject if it differs), instead of reconstructing internal field processors from wire-selected factory names (site 2).
- Reject any
PlaceholderRangeInfowithlength <= 0(or out-of-rangeoffset) with a request-scoped 4xx, and convert the encoder-runner invariant into a checked, request-scoped error rather than a process-fatalassert(site 3). - Recompute or verify the content hash for submitted
kwargs_databefore using it as a cache key, and namespace receiver-cache keys to a server-generated or principal scope, refusing cache-reads for hashes the caller did not legitimately produce (site 4). - Serialize
is_embedinPlaceholderRangeInfo, validate its length against the placeholder span, and reconstruct it on replay (site 5).
Site 1 — validate grid geometry before the vision encoder. Replace the bare assert grid_thw.ndim == 2 in _process_image_input()/_process_video_input() with a shared helper that recomputes the split sizes and rejects a grid whose patch-row count does not match the pixel tensor (and rejects non-positive / non-merge-divisible dims), so the mismatch never reaches image_embeds.split():
# vllm/model_executor/models/qwen2_vl.py — _process_image_input()
- grid_thw = image_input["image_grid_thw"]
- assert grid_thw.ndim == 2
+ grid_thw = image_input["image_grid_thw"]
+ input_type = image_input["type"]
+ input_tensor = (
+ image_input["image_embeds"]
+ if input_type == "image_embeds"
+ else image_input["pixel_values"]
+ )
+ sizes = _validate_qwen2_vl_input_geometry(
+ modality="image",
+ input_type=input_type,
+ input_tensor=input_tensor,
+ grid_thw=grid_thw,
+ spatial_merge_size=self.visual.spatial_merge_size,
+ )
...
- # Split concatenated embeddings for each image item.
- merge_size = self.visual.spatial_merge_size
- sizes = (grid_thw.prod(-1) // merge_size // merge_size).tolist()
return image_embeds.split(sizes)
where the helper raises before the encoder runs:
# vllm/model_executor/models/qwen2_vl.py — new _validate_qwen2_vl_input_geometry()
if t <= 0 or h <= 0 or w <= 0:
raise ValueError(f"{modality} grid_thw row {index} must be positive ...")
if h % spatial_merge_size != 0 or w % spatial_merge_size != 0:
raise ValueError(f"{modality} grid_thw row {index} must be divisible ...")
...
if actual_rows != expected_rows:
raise ValueError(
f"{modality} {row_kind} do not match grid_thw: "
f"expected {expected_rows}, got {actual_rows}."
)
Site 3 — constrain the placeholder schema. Make PlaceholderRangeInfo reject non-positive lengths and negative offsets at the Pydantic boundary (plus parallel-length, non-overlapping, and within-prompt validators), turning the process-fatal assert into a request-scoped 422:
# vllm/entrypoints/scale_out/token_in_token_out/protocol.py — PlaceholderRangeInfo
- offset: int
- length: int
+ offset: int = Field(ge=0)
+ length: int = Field(gt=0)
Site 4 — bind the cache key to the payload. Derive each cache key by hashing the submitted serialized tensor (ignoring the caller's mm_hashes) and refuse cache-only reads, so a forged hash can neither poison nor read a victim's slot:
# vllm/entrypoints/scale_out/token_in_token_out/serving.py — serve_tokens()
+ mm_hashes = _bind_mm_hashes_to_kwargs_data(
+ features.mm_hashes, features.kwargs_data,
+ )
engine_input = mm_input(
prompt_token_ids=request.token_ids,
mm_kwargs=MultiModalKwargsItems(mm_kwargs),
- mm_hashes=features.mm_hashes,
+ mm_hashes=mm_hashes,
mm_placeholders=mm_placeholders,
cache_salt=request.cache_salt,
)
where _bind_mm_hashes_to_kwargs_data() raises on kwargs_data is None (cache-only read) and derives sha256(modality || "\0" || serialized_item) per item. Site 2 applies the same rebind-to-declared-schema pattern in mm_serde.py (passing modality + mm_processor into decode_mm_kwargs_item), and site 5 adds an is_embed field to PlaceholderRangeInfo with from_placeholder_range/to_placeholder_range helpers so the sparse mask survives render-to-generate replay. Each site fix ships with a regression test. This packet groups the five sites because they share one entry point (/inference/v1/generate + /render) and one root cause; we are happy to split it into per-component advisories (for example, engine-fatal input-validation vs. cache-key binding vs. transport-schema integrity) if the vLLM team prefers.
Credit
Reported by: Patch the Planet (Trail of Bits + OpenAI collaboration)
These vulnerabilities were discovered using GPT-5.5-Cyber as part of the Patch the Planet security initiative.
Related Vulnerabilities
Other vulnerabilities affecting the same packages