arXiv:2605.11003cs.CRcs.AI2026-05被引 2

开放世界智能体的授权与执行偏差是重大安全漏洞,需动态检测并定位根源。

The Authorization-Execution Gap Is a Major Safety and Security Problem in Open-World Agents

论文配图:The Authorization-Execution Gap Is a Major Safety and Security Problem in Open-World Agents
图 1 · 摘自论文原文
  • 识别授权执行偏差的三大结构性根源:委托不全、通道污染、组合碎片化。
  • 小偏差可导致不可逆损害,现有防御仅治标不治本。
  • 强调运行时授权完整性检查,推动过程级安全评估新范式。

本文指出,开放世界智能体中存在严重的授权-执行偏差(Authorization-Execution Gap, AEG)问题,即主体意图授权的内容与智能体实际执行行为之间的偏离。由于智能体在工具、持久状态和多智能体协作中自主运行,微小的授权偏差就可能引发难以挽回的损害。我们论证,多数智能体失败可归因于三类结构性根源:委托层面的不完整、通信层面的污染、组合层面的碎片化,且同一故障可能源于任一来源。若不识别具体根源,仅针对表象的防御无法根除隐患。因此,智能体安全应聚焦源头诊断与防御。由于这些根源在执行过程中动态产生,必须在运行时实施授权完整性检查,而非依赖一次性前置过滤或事后审计。对NeurIPS而言,相关论文不仅应报告任务成功率等结果指标,还应提供执行过程中检测、约束并归因于特定结构来源的流程证据。

原文摘要 · Abstract (English)

This position paper argues that the Authorization-Execution Gap (AEG) is a major safety and security problem in open-world agents. The AEG is the divergence between what a principal intends to authorize and what an open-world agent ultimately executes. Because such agents act autonomously across tools, persistent state, and multi-agent handoffs, even small instances of authorization divergence can cause harm that is difficult or impossible to undo. We argue that many observed agent failures can be traced to three structural sources of AEG: delegation-level incompleteness, channel-level corruption, and composition-level fragmentation. The same observed failure may arise from any of these sources. Without identifying the source, a defense targeting the symptom alone cannot address the underlying cause. Agent safety and security should therefore emphasize source-oriented diagnosis and defense. Because the structural sources of AEG arise dynamically during execution, this approach necessarily requires authorization integrity checks applied during execution, rather than relying solely on one-shot upfront filtering or post-hoc audit. For NeurIPS, the implication is that papers on open-world agents should report not only outcome-level metrics such as task success or attack resistance, but also process-level evidence showing where AEG was detected, constrained, and attributed to a structural source during execution.

智能体安全授权偏差运行时检查

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。