Preliminary research
Cache Me, Catch You: Exploiting LLM Caching Layers in vLLM, GPTCache & Friends
AI-collected research leads through 6 October 2026, including bounded month-by-month reviews of selected social and community sources from January through September. Unranked, incomplete, not community-vetted, and subject to change.
Research on LLM serving caches (vLLM, GPTCache and peers) where the layer deciding whether two requests are 'the same' is fooled. All three cache types rest on serialize-key-reuse, so an attacker crafts colliding inputs: hash-colliding padding poisons a shared system-prompt or block-wise KV entry so a malicious block is skipped; near-neighbour embeddings make semantic and RAG caches serve a planted answer; byte-only image hashing makes moderation reuse a benign verdict. Three CVEs resulted.
Record
- Researcher
- Xiangfan Wu, Lingyun Ying, Haipeng Qu, Guoqiang Chen and Yacong Gu
- Format
- Whitepaper
In the archive
Related sources
- Paper (NDSS 2026)
- Code
- Kv Cache Collision advisory
- Image Hash Collision (tobytes) advisory
- PNG tRNS Transparency Bypass advisory
Tags
This page is the archive's own catalogue record. The research is the work of Xiangfan Wu, Lingyun Ying, Haipeng Qu, Guoqiang Chen and Yacong Gu, first published at the original source. Preserved copies are kept so the citation survives its host.