A single bad character can crash a vLLM server, and the fix is still pending

A moderate-severity flaw in the popular AI serving engine lets any logged-in user kill the whole process with one malformed request.

ThreatVectr NewsdeskAI-assistedPublished · Editor: Lee Brown· 3 min read
Illustration: a dark server rack with a single blade pulled halfway out, cooling fans still spinning
Illustration made with AI. Not a photograph of the events described.
Share

Key points

  • A flaw tracked as GHSA-2823-qmq8-rwvj lets any authenticated user crash a vLLM server by sending a single character the system should have rejected earlier.
  • The bug sits in vLLM, open-source software that runs large language models (the technology behind chatbots) for other applications to call.
  • It carries a CVSS score of 6.5 and is confirmed in vLLM up to and including version 0.25.1, with no patched release listed in the advisory at time of writing.
  • The crash only fires when the LMCache-MP connector, an add-on that speeds up responses by caching data, is turned on.
  • A separate, more serious unauthenticated code-execution bug in LMCache itself was reported this week by The Hacker News and remains unpatched.

vLLM has a validation problem, and it is the kind that gives operations teams a bad afternoon.

The project's maintainers published an advisory, GHSA-2823-qmq8-rwvj, covering a denial-of-service bug in the way the server handles a field called cache_salt. That field is a small piece of text a client can attach to a request to keep its cached data separate from someone else's. vLLM accepts it on its OpenAI-compatible endpoints, the ones that mimic the API shape most AI tools already speak.

The server checks that cache_salt is a non-empty string. That is the entire check. No length limit, no banned characters.

The value then travels, untouched, into a second library called LMCache, which does apply stricter rules. LMCache refuses any salt containing @, /, \ or a null byte, or anything longer than 128 characters. When it sees one, it raises an error.

Nobody catches that error. The scheduling code that called into LMCache has no safety net around it, so the error bubbles up and takes the whole serving process down with it. One request, one dead server.

Who is actually at risk?

Operators running vLLM 0.25.1 or earlier with the built-in LMCache-MP KV connector switched on. If that connector is not enabled, the bad path is never reached. The attacker does need valid credentials to the API (CVSS vector PR:L), so this is not an open-internet drive-by. In practice, that still includes any multi-tenant setup where several teams or customers share one vLLM instance behind a gateway.

The impact is availability only. No data is read, no data is changed. The process simply exits.

What should operators do right now?

Turn off the LMCache-MP connector if you can live without it, or put an input filter in front of vLLM that rejects any cache_salt containing @, /, \, a null byte, or more than 128 characters. That matches what LMCache enforces downstream, so sanitising at the edge removes the trigger.

Here are the facts in one place.

Item Value
Advisory ID GHSA-2823-qmq8-rwvj
CVSS v3.1 6.5 (Moderate)
CWE CWE-20, CWE-248
Affected vLLM ≤ 0.25.1 (confirmed on commit 752a3a504485)
Trigger LMCache-MP KV connector enabled
Impact Server process crash (DoS)

How does this fit with the LMCache code-execution bug?

It rhymes with it, and that is the part worth watching. The Hacker News reported this week on a separate, unauthenticated remote code execution flaw in LMCache's multiprocess mode, where the cache runs as its own server reachable over the ZeroMQ messaging library. That one has no fix either.

Both bugs live at the seam between vLLM and LMCache, where one project assumes the other is doing the checking. My read: the AI serving stack has grown faster than its input-validation discipline, and the pattern of trusting a sibling library to sanitise your inputs is going to produce more of these before it produces fewer. If you run vLLM in production, inventory every optional connector you have enabled and ask what crashes when the thing on the other end says no.

© 2026 Threat Vectr