Web Hack List

Preliminary research

Prompt Injection as Role Confusion (CoT Forgery)

AI-collected research leads through 6 October 2026, including bounded month-by-month reviews of selected social and community sources from January through September. Unranked, incomplete, not community-vetted, and subject to change.

Language models receive system, user, tool and reasoning content as one token stream distinguished only by role tags. Using linear probes trained on identical text wrapped in each tag, this work shows models infer role largely from writing style, which overrides the true tag, and that the measured role attribution of injected text predicts attack success. It recasts prompt injection as role confusion and demonstrates forged-reasoning and role-claiming attacks that follow.

Record

Researcher
Charles Ye, Jasmine Cui and Dylan Hadfield-Menell

In the archive

Related sources

Tags

This page is the archive's own catalogue record. The research is the work of Charles Ye, Jasmine Cui and Dylan Hadfield-Menell, first published at the original source. Preserved copies are kept so the citation survives its host.