CVE-2026-44223
vLLM is an inference and serving engine for large language models (LLMs). From to before 0.20.0, the extract_hidden_states speculative decoding proposer in vLLM returns a tensor with an incorrect shape after the first decode step, causing a RuntimeError that crashes the EngineCore process. The crash is triggered when any request in the batch uses sampling penalty parameters (repetition_penalty, frequency_penalty, or presence_penalty). A single request with a penalty parameter (e.g., "repetition_penalty": 1.1) is sufficient to crash the server. This vulnerability is fixed in 0.20.0.
Scoring
- Severity
- MEDIUM
- CVSS base score
- 6.5
- CVSS vector
- CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H
- EPSS probability
- 0.37%
- CWE
- CWE-131, CWE-704
- Published
- 2026-05-12
- Last modified
- 2026-06-22
Affected products
- vllm-project vllm
Weakness type
Related vulnerabilities
- CVE-2026-22590 — Fast-DDS Discovery Server: Out-of-Bounds Read & Heap Memory Disclosure via DATA_FRAG sampleSize / fragmentsInSubmessage
- CVE-2026-69598 — Windows iSCSI Remote Code Execution Vulnerability
- CVE-2026-78221 — An incorrect buffer size calculation in the Windows Interactive Service in OpenVPN 2.7_alpha1...
- CVE-2026-18743 — Popt-devel: popt-static: short realloc in poptconfigfiletostring
- CVE-2026-78002 — Rsyslog: rsyslog: denial of service via heap buffer overflow in rainerscript replace() function
- CVE-2026-44254 — Wazuh: Stack Out-of-Bounds Write in remoted Decompression Path
- CVE-2026-52834 — jxl-oxide: Out-of-bounds writes due to integer overflow in jxl-grid on 32-bit platforms
- CVE-2026-75093 — sonos tract ONNX Initializer Loader tensor.rs from_raw_dt_align buffer size