arXiv:2602.05066cs.CRcs.AI2026-02被引 1

AI代理被当通信工具,绕过安全监控完成攻击

Bypassing AI Control Protocols via Agent-as-a-Proxy Attacks

  • 用代理充当攻击信使,同时避开代理和监控系统
  • 在AgentDojo上对多种监控模型攻击成功率高
  • 大模型监控也易被同等级代理突破,防御很脆弱

随着AI代理自动化关键任务,它们仍易受间接提示注入(IPI)攻击。现有防御依赖监控协议,联合评估代理的思维链(CoT)与工具使用行为以确保与用户意图对齐。我们展示了一种新型「代理作代理」攻击,攻击者将代理作为传递媒介,同时绕过代理与监控系统。尽管先前研究关注小监控模型能否监督大代理,我们发现即使是前沿规模的监控模型(如Qwen2.5-72B)也仍可被具备相似能力的代理(如GPT-4o mini、Llama-3.1-70B)攻破。在AgentDojo基准测试中,针对AlignmentCheck和Extract-and-Evaluate监控器,在多种监控大模型下均实现高攻击成功率。结果表明,当前基于监控的智能体防御机制无论模型规模大小都存在根本性脆弱性。

原文摘要 · Abstract (English)

As AI agents automate critical workloads, they remain vulnerable to indirect prompt injection (IPI) attacks. Current defenses rely on monitoring protocols that jointly evaluate an agent's Chain-of-Thought (CoT) and tool-use actions to ensure alignment with user intent. We demonstrate that these monitoring-based defenses can be bypassed via a novel Agent-as-a-Proxy attack, where prompt injection attacks treat the agent as a delivery mechanism, bypassing both agent and monitor simultaneously. While prior work on scalable oversight has focused on whether small monitors can supervise large agents, we show that even frontier-scale monitors are vulnerable. Large-scale monitoring models like Qwen2.5-72B can be bypassed by agents with similar capabilities, such as GPT-4o mini and Llama-3.1-70B. On the AgentDojo benchmark, we achieve a high attack success rate against AlignmentCheck and Extract-and-Evaluate monitors under diverse monitoring LLMs. Our findings suggest current monitoring-based agentic defenses are fundamentally fragile regardless of model scale.

AI安全提示注入智能体防御

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。