
LMCache's Unpatched RCE: Reproducing the Public PoC Against an LLM Cache
LMCache Has a Public RCE PoC and No Patch
This post covers what is actually confirmed about the unpatched LMCache RCE, how to stand up a disposable lab to reproduce the public PoC without touching a real inference cluster, what the path from cache port to code execution looks like, and what to detect and block while no fix exists.
What the October 2026 reporting actually says
On 2026-10-08, CyberSecurityNews ran "PoC Released for Critical LMCache Flaw Enabling Unauthenticated Remote Code Execution," and GBHackers published a companion piece the same day: the flaw "remains unpatched" and a public exploit is out. That is the whole confirmed surface. A working proof-of-concept exists, it is described as unauthenticated remote code execution, and there was no fix when those pieces published. No CVE identifier, no affected version range, and no vendor advisory appear in the material I have. I have not reproduced the exploit.
Where I land: unauthenticated RCE in an LLM KV cache tier is worse than the same bug in a general-purpose cache. Look at what sits next to the cache process — provider API keys, object-store credentials, model weights, and often the same node holding the GPUs. A Redis with a bad deserializer is bad. A KV cache with one is a credential-theft primitive and a cross-tenant data leak rolled together. That is a keep-you-up-at-night bug, not a next-sprint bug.
Why an LLM KV cache is not just another cache
LMCache sits between a vLLM-style inference worker and a shared cache or storage backend. The worker computes key/value tensors during prefill; LMCache offloads them so a later request sharing a prefix skips recomputation. Serialized tensor payloads therefore move over the network between the serving process and a cache server, and on to Redis, local disk, or object storage. It is on the critical path, it holds data derived from prompts, and it speaks a protocol that has to round-trip arbitrary structured blobs. That is why the deserialization surface matters.
What the public LMCache sources confirm
- A public proof-of-concept exists for the LMCache flaw.
- The flaw is described as unauthenticated remote code execution.
- As of the reporting date, no patch was available.
What the public sources do not confirm
- No CVE identifier, affected version range, or vendor security advisory appears in the material I have.
- The exact vulnerable component, endpoint, and payload format are not established.
- Whether a default deployment configuration exposes the vulnerable path is unknown to me.
Everything below is inference to check against the PoC and the project source, not quoted advisory fact.
Standing Up a Disposable Lab First
I would not run a public RCE PoC anywhere that can reach a real inference cluster, a GPU node, or a cloud metadata endpoint. Those three are exactly what an unauthenticated RCE in a cache tier is looking for, and a lab that can reach them is not a lab.
# Pin by digest. "latest" in a security test is how you lose track of what you tested.
export LMCACHE_IMAGE='lmcache/vllm-openai@sha256:<digest>'
docker network create lmcache-lab-net
docker run -d --name lab-cache --network lmcache-lab-net --read-only --tmpfs /tmp:rw,noexec,nosuid,size=64m --cap-drop ALL --security-opt no-new-privileges --pids-limit 256 --memory 4g "$LMCACHE_IMAGE"
## Driver container: same network, no host network, no inherited cloud env.
docker run --rm -it --network lmcache-lab-net --env-file /dev/null --cap-drop ALL alpine:3.20 shA --internal network removes egress entirely, which is stronger — but it also breaks host port publishing, so you drive the test with docker exec instead of from the host shell. Take that trade.
Keep the test host off the corporate network and off any machine that holds IAM credentials. A cache-tier RCE is a credential-theft primitive first and a code-execution bug second.
Tracing the Attack Path From Cache Port to Code Execution
The exposed cache service and its reachable surface
Enumerate what the process actually listens on rather than trusting a default port number from a blog post. On a serving host:
# From the host
sudo ss -ltnp | grep -i -e lmcache -e vllm -e python
## From inside the container (no ss? use /proc)
docker exec lab-cache sh -c 'for p in /proc/[0-9]*; do ls -l $p/exe 2>/dev/null; done | sort -u'Record which ports the cache server and any controller process bind, then check who can reach them. In my experience these listeners land in container or VM networks configured flat because "it's all internal" — which is the same as saying they are reachable from every workload that gets compromised anywhere on that network.
The likely injection point: deserializing cache payloads
Marked inference: KV cache payloads are serialized structured tensors, and in Python the obvious dangerous primitive is a deserializer that trusts its input — pickle, or torch.load without weights_only=True. A remote cache backend that receives a blob and reconstructs a Python object from it is an RCE by design if that blob is unauthenticated. Which deserializer LMCache uses, and on which path, I have not confirmed. Neither the wire format. Verify against the source before you write a report.
A minimal verification transcript
The honest way to produce evidence here is to capture the PoC's own traffic on the wire and record it, then confirm the impact inside the container. I did not run this, so the fields below are placeholders — fill them from your run rather than quoting mine.
CACHE_PORT=$(docker exec lab-cache sh -c "awk '/: [0-9A-F]*:/ {split($2,a,":"); print strtonum("0x"a[2])}' /proc/net/tcp" | head -1)
## Second terminal: capture the request the PoC sends
docker run --rm --network lmcache-lab-net --cap-add NET_RAW nicolaka/netshoot tcpdump -i any -A -s0 -w /tmp/poc.pcap "tcp port $CACHE_PORT"
## After the run
docker run --rm -v /tmp:/tmp nicolaka/netshoot tcpdump -r /tmp/poc.pcap -A 'tcp[((tcp[12:1] & 0xf0) >> 2):4] = 0x504f5354'| Field to record | Placeholder | Why it matters |
|---|---|---|
| Request line and path | <method> <path> | Proves reachability without credentials |
| Auth headers present | <none / bearer / hmac> | The "unauthenticated" claim lives here |
| Content-Type | <value> | Tells you it is a serialized blob, not JSON |
| Response status | <code> | 200 means the server processed attacker data |
| Response body | <body> | Confirms or refutes code execution on the host |
Then prove impact with output, not adjectives. Inside the container, after the PoC fires, capture this verbatim:
docker exec lab-cache sh -c 'id; hostname; cat /proc/1/cgroup; env | cut -d= -f1 | sort'If id no longer prints the container's unprivileged user, or env names a provider key, the impact is demonstrated. If it prints exactly what it printed before the PoC, you have not reproduced the bug — say so.
Impact: What RCE in the Cache Tier Actually Buys
| Layer | What the attacker reaches | Why it matters |
|---|---|---|
| Cache process | Host filesystem, GPU node, other pods on the node | The cache is usually co-located with inference, not isolated from it |
| Environment | Model provider keys, object-store keys, HF tokens | Converts a compute compromise into a data and billing compromise |
| Shared cache | Other tenants' prompt and response data | KV blocks are derived from real user input |
| Flat cluster network | The inference API behind the cache | The cache becomes a pivot with valid internal addressing |
Ranked by severity, credential exposure goes first, not code execution. Pull provider and object-store keys out of the cache pod and you cap the blast radius even if the RCE still fires — the attacker goes from "spend your budget and read your storage" to "shell in a box that talks to nothing." But the fix I would actually ship first is reachability: if the listener has no authentication, identity-gated network access is the only control that removes the precondition the whole bug depends on. Everything else is damage limitation.
Detection When You Cannot Patch LMCache
Detection is unusually cheap right now, and that is the argument for doing it today. Almost nothing legitimate talks to a KV cache port except the worker set you already know. The baseline is nearly empty, so the signal-to-noise ratio is close to one.
| Signal | Check | Legitimate baseline |
|---|---|---|
| Unexpected clients to the cache port | Connection table filtered to that port, excluding known worker IPs | Only your vLLM workers |
| Pickle-framed blobs from non-worker sources | Payload capture on that port | None — workers only |
| Child processes spawned by the cache process | Process tree of the cache PID | None, or a fixed supervisor set |
| First-seen outbound connections | Egress from the cache pod to anything but your storage backend | Storage backend only |
CACHE_PORT=65432 # set to the port you recorded from ss
sudo ss -tnp state established "( dport = :$CACHE_PORT or sport = :$CACHE_PORT )" | awk 'NR>1 {print $5, $6}' | sort | uniq -c | sort -rnRun that on the serving host at a quiet hour, save the output, and alert on anything that appears later. Column numbering shifts between ss versions — verify against your own output once. If you have process-level telemetry, the child-process check is the highest-fidelity single signal: a cache server that never forks should never fork.
Mitigation for an Unpatched LMCache
Network isolation buys time. It is not a fix, and "restrict to internal" fails the moment anything else on the internal network is compromised — which is the premise of the attack path above.
| Control | What it stops | What it does not stop |
|---|---|---|
| Deny ingress to the cache port from everything but known worker identities | The unauthenticated precondition | A compromised worker, or a spoofed identity on a flat network |
| Run the cache process as non-root with a read-only root filesystem | Trivial persistence and some file writes | Reading mounted secrets and env vars |
| Remove cloud and model-provider credentials from the pod | Credential theft via env or filesystem | Lateral movement inside the cluster |
| Drop the remote cache backend for single-node deployments | The network-reachable deserializer entirely | Nothing — this is a real reduction in exposure |
| Rotate every secret that lived in the pod | Retroactive use of stolen keys | Any exfiltration that already happened |
On isolated single-node setups, kill the remote cache backend. The performance loss is measurable and you can put a number on it; the reduction in exposure is binary.
What I Confirmed vs What I Did Not Test
Confirmed from primary reporting: a public proof-of-concept exists for the LMCache flaw; it is described as unauthenticated remote code execution; no patch was available at time of writing, per the 2026-10-08 GBHackers report.
Not tested by me: the exact vulnerable code path; whether a default deployment configuration exposes it; whether any partial fix or advisory has landed since publication; the endpoint and payload format. The deserialization hypothesis is inference, not a finding. Keep that separation visible in internal write-ups — an unverified root cause burns credibility fast.
Further Reading
- PoC Released for Critical LMCache Flaw Enabling Unauthenticated Remote Code Execution — CyberSecurityNews, 2026-10-08.
- Critical LMCache RCE Vulnerability Remains Unpatched, Public PoC Exploit Available — GBHackers News, 2026-10-08.
- LMCache repository — check the commits and releases yourself rather than trusting a news cycle.
- LMCache issue tracker — the place a real advisory or fix would land first.


