arXiv:2608.12789cs.CRcs.AI2026-08

用来源溯源和语义先验防御工具代理的环境欺骗攻击

PIPES: Securing Agent Perception with Provenance and Priors

论文配图:PIPES: Securing Agent Perception with Provenance and Priors
图 1 · 摘自论文原文
  • 通过内容来源和语义预期双重验证筛选响应单元
  • 对抗攻击成功率从84.7%降至2.3%,良性任务保持92.5%效率
  • 适合关注AI代理安全与可信推理的研究者

使用工具的智能体依赖外部数据,但其响应极少标注来源或语义意图。我们揭示此缺陷可被用于状态污染攻击:攻击者伪造内容使代理误信超出其信息权限的环境事实,导致行为看似合理。为此提出PIPS(溯源感知、先验约束筛选),基于静态字段契约或可信来源元数据,对响应单元进行语义先验与来源层级验证。违反规则的内容会被标记,部署系统可选择移除、警告、阻断或升级。在Gemma 4 31B IT作为目标代理的VitaBench与AgentDyn三组测试中,针对自适应PAIR风格攻击,平均攻击成功率由84.7%降至2.3%,同时维持92.5%的良性任务性能(无防御时为90.6%)。

原文摘要 · Abstract (English)

Tool-using agents consume external data from sources with different levels of trust, yet tool responses rarely identify who produced each component or what it should convey. We show that this gap enables state-corruption attacks, in which attacker-controlled content makes environmental claims beyond the informational authority of its response component and corrupts the agent's perceived environment, making the resulting action appear justified to existing guardrails. We introduce PIPES (Provenance-Informed, Prior-Enforced Screening), which screens response units using semantic priors and source provenance. PIPES uses static field contracts when schemas provide stable expectations, and conditions screening of open-ended content on the pre-response trajectory and trusted provenance metadata. It marks units that violate their semantic prior or the provenance hierarchy; deployments may remove, warn, block, or escalate detected violations. We instantiate atomic removal and evaluate PIPES against adaptive PAIR-style attacks. Across the three VitaBench and three AgentDyn splits with Gemma 4 31B IT as the target agent, PIPES reduces average attack success from 84.7% to 2.3%, while preserving average benign utility (92.5% with PIPES versus 90.6% without defense).

AI安全代理系统溯源验证对抗攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。