A single message, running as root
JFrog’s security research team disclosed CVE-2026-105192 on 7 October, rating it 9.8 and critical. It was found by Yuval Moravchick of the same team. As of the advisory there is no fixed version.
LMCache is the key-value cache layer that vLLM deployments use to keep attention caches alive across requests and machines. In multiprocess mode it opens a ZeroMQ ROUTER socket with no CURVE encryption, no ZAP authentication, no password and no message authentication at all. Messages arrive as msgpack. Extension code 1 is registered for a type called DeviceIPCWrapper, whose Deserialize method calls Python’s pickle.loads.
That call happens while the server is decoding the arguments to a REGISTER_KV_CACHE request — before any handler logic runs. So a single unauthenticated message executes code as the LMCache process user. The official container images run that process as root.
Whether you are exposed comes down to one flag
The 9.8 is not the whole story, and JFrog says so. The transport binds to localhost unless an operator sets a routable address with --host, and the critical score applies to that routable configuration. In the advisory’s words, “A stock single-host install that leaves the default bind is not reachable from other machines.”
LMCache used only inside a vLLM process does not open the port at all. The exposed shape is the multi-node deployment — the one you reach for when a single box stops being enough, which is also the one most likely to be running production traffic.

No patch, so the mitigations are yours
Versions 0.3.9 and later are affected. JFrog’s asks of the project are to replace the serializer behind msgpack extension code 1 with a safe format, to stop calling pickle.loads on data from unauthenticated peers, to authenticate the transport with CURVE or a per-message HMAC, and to refuse a routable bind unless authentication is configured.
Until a fixed release exists, the advice to operators is blunt: do not set --host to a routable address, keep port 5555 on localhost or a trusted cluster network, and firewall it. JFrog adds the caveat that matters — anyone who can connect can still run code as the LMCache user.
The same bug class, twice in a week
Unpickling data that came off the network is a Python mistake with a twenty-year paper trail, and it is now turning up in the plumbing of model serving. At Pwn2Own Ireland the same week, LiteLLM fell on the first day to improper input validation chained with code injection, worth $40,000, and a second team took it with four more bugs.
The AI-specific part of this is not the vulnerability. It is that a cache layer sitting between a model server and its GPUs was built with an unauthenticated control plane, because the deployment was assumed to live on a trusted network.

What to check tonight
Two questions settle most of it: is port 5555 reachable from anywhere but the host, and does the container run as root. If the answer to both is yes, the advisory describes a path from a single packet to a shell on a machine with your model weights on it.
The thing to watch is the fixed release. Until then this is a known, scored, publicly documented hole with a patch gap, in a component that most teams running vLLM at scale did not choose deliberately so much as inherit.