What happened
NVD published a nine-CVE batch on 2026-09-26 covering vLLM through 0.29.0. Most are unauthenticated denial-of-service/resource-exhaustion flaws (CVSS 5.3–6.5): oversized/multimodal payloads consume memory/CPU during decoding or crash the EngineCore/decode worker requiring service restart; one (CVE-2026-100653) is a model-revision supply-chain pinning gap. Fixes are in 0.29.0.
Why it matters
These are attacker-reachable DoS primitives against the core inference tier: a single crafted multimodal or stop-token request can take down a shared GPU serving endpoint (and in disaggregated serving, poison the decode worker shared by many tenants), which is a direct availability hit to AI deployments serving public traffic; the revision-pin gap is a supply-chain hardening gap for multi-modal model loading.
Attack vector
Multiple paths: media fetched/materialized before the VLLM_MAX_AUDIO_CLIP_FILESIZE_MB cap is enforced (CVE-2026-100650/100648), missing audio decode size limit (CVE-2026-100648), out-of-vocabulary stop_token_ids reaching CUDA index_put_ and crashing EngineCore when min_tokens>0 (CVE-2026-100652/100654), missing decoder prompt-length validation on the disaggregated /inference/v1/generate endpoint (CVE-2026-100651), unvalidated long cache_salt processed on the single scheduler thread (CVE-2026-100647), HF revision pin not propagated for FunAudioChat/Tarsier2 artifact loads enabling supply-chain drift (CVE-2026-100653), and PyNvVideoCodec decoder-limit bypass (CVE-2026-100649).
Mitigation
Upgrade to vLLM 0.29.0 or later; pin --revision / --code-revision for FunAudioChat/Tarsier2 models; enforce size caps and per-modality limits at the ingress/API gateway in front of the engine for unpatched versions.