Web Hack List

Preliminary research

Assessing Automated Prompt Injection Attacks in Agentic Environments

AI-collected research leads through 6 October 2026, including bounded month-by-month reviews of selected social and community sources from January through September. Unranked, incomplete, not community-vetted, and subject to change.

The paper adapts white-box GCG and black-box TAP attacks to prompt injection against agents in AgentDojo, evaluating 80 task pairs across four domains and multiple models. Black-box optimization performs better under the tested budgets, while transfer to frontier models remains limited and model-dependent.

Record

Researcher
David Hofer, Edoardo Debenedetti and Florian Tramèr
Published by
arXiv.org

In the archive

Tags

This page is the archive's own catalogue record. The research is the work of David Hofer, Edoardo Debenedetti and Florian Tramèr, first published at the original source. Preserved copies are kept so the citation survives its host.