Request-selected PyNvVideoCodec GPU decode bypasses static VRAM reservation
vLLM security advisories
All 67 advisories vLLM has published with an identifier, newest first. Severity is the one its publisher assigned, and the fix is the release the publisher named. Nothing on this page is our judgement.
- Advisories
- 6761 carry a CVE
- critical
- 5
- high
- 18
- medium
- 41
- low
- 3
- Fix in the archive
- 67of 67 matched to a release
- Oldest
- 27 Jan 20251.6 years ago
Every advisory here is matched to the release in the archive that carries its fix. This page is a copy of what the publisher published, kept for reference. The authoritative source for a security question is the publisher, and an advisory missing from here is not evidence that none exists. What this page does and does not tell you sets out the limits in full.
Newest first
Every productSpeech-to-text audio decode duration limit bypass via forged header sample rate
Denial of Service NanoNemoTronVL Video Audio Extraction Bomb
LlavaOnevision2 processor loader executes attacker model code with `trust_remote_code=False` (inert `trust_remote_code` kwarg to `transformers.get_class_from_dynamic_module`) — RCE from a malicious model
Unauthenticated audio decompression-bomb DoS in /v1/chat/completions: VLLM_MAX_AUDIO_DECODE_DURATION_S guard not wired into the chat audio path (sibling of CVE-2026-5497)
SSRF + arbitrary local file read in MiMoV2OmniMultiModalProcessor `_fetch_image` and audio loader bypass MediaConnector protections
Cross-User Data Leak Vulnerability
vLLM Unauthenticated Requests Exploit DeepStream Backend Confusion for DoS
Unauthenticated Internal Path and Username Disclosure via Validation Error Messages
Derender endpoints decode caller-supplied GenerateResponse token IDs without output bounds
ReDoS via structured_outputs.regex in the lm-format-enforcer backend (no compile timeout) — missed sibling of GHSA-rwxx-mrjm-wc2m
Completion prompt lists fan out into unbounded engine requests
Incomplete CVE-2025-62164 remediation can be bypassed by concurrent prompt parts
ReDoS via structured_outputs.regex compiled without timeout in xgrammar and outlines backends
Remote DoS in vLLM via Invalid Recovered Token Reinjection
DoS caused by sending `/v1/completions` with prompt embeds payload with models that use M-RoPE
Speech-to-text upload size limit is enforced after full UploadFile read
Security Check Bypass via assert Statement in Activation Function Loading Allows Arbitrary Code Execution
vLLM image EXIF Rotation & PNG tRNS Transparency Not Normalized, Causing Mismatch Between Model Input and Expectations
GGUF dequantize kernel int truncation exposes uninitialized GPU memory in multi-tenant serving
OOM Denial of Service via Audio Decompression Bomb
temperature=NaN and temperature=Infinity bypass validation and propagate to GPU kernels
vLLM: incomplete CVE-2026-22778 fix leaks PIL repr addresses via Anthropic router (CWE-532)
Artifact Pin Decay in vLLM allows pinned deployments to load unpinned code, weights, and processors
Dependency Confusion Vulnerability in vLLM Dockerfile
OpenAI API Auth Bypass
extract_hidden_states speculative decoding crashes server on any request with penalty parameters
Remote DoS via Special-Token Placeholders
OOM Denial of Service via Unbounded `n` Parameter in OpenAI API Server
Server-Side Request Forgery (SSRF) in `download_bytes_from_url `
Denial of Service via Unbounded Frame Count in video/jpeg Base64 Processing
Downmix Implementation Differences as Attack Vectors Against Audio AI Models
Hardcoded trust_remote_code=True in NemotronVL and KimiK25 bypasses user security opt-out
SSRF Protection Bypass in vLLM
vLLM RCE In Video Processing
Server-Side Request Forgery (SSRF) in `MediaConnector`
DoS via incorrect shape of multimodal embedding inputs
RCE via auto_map dynamic module loading during model initialization
DoS in Idefics3 vision models via image payload with ambiguous dimensions
Missing validation of multimodal embeddings leading to DoS and potential RCE
Remote code execution via transformers_utils/get_config
DoS via large Chat Completion or Tokenization requests with specially crafted `chat_template_kwargs`
prompt_embeds deserialization allows DoS and potential RCE
DoS with incorrect shape of multimodal embedding inputs
Server-Side Request Forgery (SSRF) in `MediaConnector`
Resource-Exhaustion (DoS) through chat_template / chat_template_kwargs in OpenAI-Compatible Server
API key authentication vulnerable to timing attack
HTTP header size limits not enforced, allowing DoS from unauthenticated requests
Remote code execution in the vllm tool call parser for Qwen3-Coder
Potential Timing Side-Channel Vulnerability in vLLM’s Chunk-Based Prefix Caching
DOS: Remotely kill vllm over http with invalid JSON schema
Regular Expression Denial of Service (ReDoS, Exponential Complexity) Vulnerability in `pythonic_tool_parser.py`
Weakness in MultiModalHasher Image/video Hashing Implementation in vLLM
clients can crash the openai server with invalid regex
A series of simple Redos in vllm.
DoS via Malformed pattern and type Fields in vLLM Tool Schema
Remote Code Execution via PyNcclPipe Communication Service
Remote Code Execution Vulnerability in vLLM Multi-Node Cluster Configuration
Denial of Service via ZeroMQ on Multi-node vLLM Deployment
Remote Code Execution via Mooncake Integration
phi4mm: Quadratic Time Complexity in Input Token Processing leads to denial of service
CVE-2025-24357 Malicious model remote code execution fix bypass with PyTorch < 2.6.0
Denial of Service by abusing xgrammar cache
Remote Code Execution via Mooncake Integration
Denial of Service by abusing outlines unbounded cache on disk
vLLM using built-in hash() from Python 3.12 leads to predictable hash collisions in vLLM prefix cache
Malicious model to RCE by torch.load in hf_model_weights_iterator