Preliminary research
Prompt Injection as Role Confusion (CoT Forgery)
AI-collected research leads through 6 October 2026, including bounded month-by-month reviews of selected social and community sources from January through September. Unranked, incomplete, not community-vetted, and subject to change.
Language models receive system, user, tool and reasoning content as one token stream distinguished only by role tags. Using linear probes trained on identical text wrapped in each tag, this work shows models infer role largely from writing style, which overrides the true tag, and that the measured role attribution of injected text predicts attack success. It recasts prompt injection as role confusion and demonstrates forged-reasoning and role-claiming attacks that follow.
Record
- Researcher
- Charles Ye, Jasmine Cui and Dylan Hadfield-Menell
In the archive
Related sources
Tags
This page is the archive's own catalogue record. The research is the work of Charles Ye, Jasmine Cui and Dylan Hadfield-Menell, first published at the original source. Preserved copies are kept so the citation survives its host.