自动挖掘大模型代理中的隐蔽提示注入漏洞
PI-Hunter: Automated Red-Teaming for Exposing and Localizing Prompt Injections

- 构建带源感知的测试用例,通过反馈迭代诱导代理暴露隐藏指令
- 在多个基准上显著提升漏洞发现率和攻击面覆盖范围
- 适合安全研究人员与模型开发者用于主动检测提示注入风险
大型语言模型正快速演变为与外部工具和环境交互的智能体,带来通过不可信外部源引发的间接提示注入攻击等新安全风险。现有防御主要聚焦推理时的内容拦截,而当前红队测试方法多以提高攻击成功率为目标,导致开发者难以看清潜在提示注入如何在智能体中产生与传播。本文提出 PI-Hunter,一种自动化智能体审计框架,用于主动暴露大模型代理中的安全漏洞。PI-Hunter 构建真实场景下的源感知测试用例,并通过反馈驱动的探索过程,促使代理从外部环境中检索并暴露嵌入的恶意指令。在多个基准、代理架构、攻击类型与防御机制上的广泛实验表明,PI-Hunter 在漏洞暴露和攻击面覆盖方面显著优于强基准自动化红队方法,且在现有提示注入防御下仍具有效性。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are rapidly evolving into agentic systems that interact with external tools and environments, introducing new security risks such as indirect prompt injection attacks through untrusted external sources. Existing defenses mainly focus on blocking malicious content at inference time, and current red-teaming methods primarily optimize attack success. As a result, developers have limited visibility into how latent prompt injections emerge and propagate through agents. We propose PI-Hunter, an automated agentic auditing framework for proactive vulnerability exposure in LLM agents. PI-Hunter constructs realistic source-aware test cases and iteratively evolves them through feedback-driven exploration to induce agents to retrieve and reveal latent malicious instructions embedded within external environments. Extensive experiments across multiple benchmarks, agent architectures, attacks, and defenses demonstrate that PI-Hunter substantially improves vulnerability exposure and attack-surface coverage over strong automated red-teaming baselines, while remaining effective under existing prompt injection defenses.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。