Web Hack List

Preliminary research

Cache Me, Catch You: Exploiting LLM Caching Layers in vLLM, GPTCache & Friends

AI-collected research leads through 6 October 2026, including bounded month-by-month reviews of selected social and community sources from January through September. Unranked, incomplete, not community-vetted, and subject to change.

Research on LLM serving caches (vLLM, GPTCache and peers) where the layer deciding whether two requests are 'the same' is fooled. All three cache types rest on serialize-key-reuse, so an attacker crafts colliding inputs: hash-colliding padding poisons a shared system-prompt or block-wise KV entry so a malicious block is skipped; near-neighbour embeddings make semantic and RAG caches serve a planted answer; byte-only image hashing makes moderation reuse a benign verdict. Three CVEs resulted.

Record

Researcher
Xiangfan Wu, Lingyun Ying, Haipeng Qu, Guoqiang Chen and Yacong Gu
Format
Whitepaper

In the archive

Related sources

Tags

This page is the archive's own catalogue record. The research is the work of Xiangfan Wu, Lingyun Ying, Haipeng Qu, Guoqiang Chen and Yacong Gu, first published at the original source. Preserved copies are kept so the citation survives its host.