arXiv:2604.18874cs.AI2026-04ACL被引 4

揭露智能体工具依赖的致命弱点:环境欺骗可致其误信假信息或陷入死循环。

How Adversarial Environments Mislead Agentic AI?

论文配图:How Adversarial Environments Mislead Agentic AI?
图 1 · 摘自论文原文
  • 提出对抗性环境注入(AEI)威胁模型,模拟工具输出被篡改的虚假场景。
  • 在5个前沿智能体上测试超1.1万次,发现抗骗能力与导航鲁棒性相互制约。
  • 开发POTEMKIN框架,支持即插即用的鲁棒性评估,适配主流智能体系统。

工具集成智能体的部署依赖外部工具提供真实输出,但这一依赖也带来关键攻击面。当前评估仅在良性环境下测试‘能否正确使用工具’,却从未考察‘若工具说谎会怎样’。我们识别出‘信任缺口’:智能体被评估的是性能,而非怀疑能力。将此漏洞形式化为对抗性环境注入(AEI),即攻击者污染工具输出以欺骗智能体。AEI构成环境欺骗——围绕智能体构建由污染搜索结果和虚构参考网络组成的‘虚假世界’。我们通过POTEMKIN实现该威胁模型,这是一个兼容模型上下文协议(MCP)的即插即用鲁棒性测试框架。识别出两种正交攻击面:‘幻象’(广度攻击)通过污染检索诱导认知漂移至错误信念;‘迷宫’(深度攻击)利用结构陷阱导致策略崩溃并陷入无限循环。在5个前沿智能体上进行超过11,000次运行测试,发现抗一类型攻击常加剧对另一类型的脆弱性,表明认知与导航鲁棒性是独立能力。

原文摘要 · Abstract (English)

Tool-integrated agents are deployed on the premise that external tools ground their outputs in reality. Yet this very reliance creates a critical attack surface. Current evaluations benchmark capability in benign settings, asking "can the agent use tools correctly" but never "what if the tools lie". We identify this Trust Gap: agents are evaluated for performance, not for skepticism. We formalize this vulnerability as Adversarial Environmental Injection (AEI), a threat model where adversaries compromise tool outputs to deceive agents. AEI constitutes environmental deception: constructing a "fake world" of poisoned search results and fabricated reference networks around unsuspecting agents. We operationalize this via POTEMKIN, a Model Context Protocol (MCP)-compatible harness for plug-and-play robustness testing. We identify two orthogonal attack surfaces: The Illusion (breadth attacks) poison retrieval to induce epistemic drift toward false beliefs, while The Maze (depth attacks) exploit structural traps to cause policy collapse into infinite loops. Across 11,000+ runs on five frontier agents, we find a stark robustness gap: resistance to one attack often increases vulnerability to the other, demonstrating that epistemic and navigational robustness are distinct capabilities.

智能体安全对抗攻击工具依赖鲁棒性评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。