CyberRota Analysis
AI-GeneratedVersions 0.22.0 through 0.23.0 of vLLM are vulnerable due to inadequate validation of stop_token_ids, permitting out-of-vocabulary token IDs to bypass checks and potentially cause CUDA tensor indexing failures. This can lead to EngineCore entering a fatal state, necessitating a service restart. Organizations utilizing these vLLM versions, particularly those relying on Rust HTTP and gRPC frontends, should prioritize patching to mitigate service disruptions.
Public Exploit Signal
A public exploit, PoC, GitHub repository or Metasploit reference was detected for this CVE.
Note: these links are listed for security research and verification purposes only.
Original NVD Description
vLLM versions 0.22.0 through 0.23.0 fail to validate stop_token_ids against vocabulary bounds in Rust HTTP and gRPC frontends, allowing out-of-vocabulary token IDs to reach MinTokensLogitsProcessor. Attackers can submit requests with min_tokens greater than zero and out-of-vocabulary stop_token_ids to trigger CUDA tensor indexing failures that leave EngineCore in a fatal state requiring service restart.
Related CVEs
Other vulnerabilities affecting the same vendor(s)