arXiv:2511.08487cs.MAcs.CL2025-11

揭示复杂任务中隐藏恶意意图的智能体安全脆弱性

How Brittle is Agent Safety? Rethinking Agent Risk under Intent Concealment and Task Complexity

  • 构建双维度评估框架,分析意图隐蔽与任务复杂度对安全的影响
  • 发现意图越隐蔽,安全对齐越差;难题反而更安全因能力受限
  • 发布OASIS基准和仿真环境,支持深度安全测试

当前大模型驱动智能体的安全评估主要聚焦于单一危害,未能涵盖恶意意图被隐藏或稀释在复杂任务中的高阶威胁。本文通过意图隐蔽性与任务复杂度两个正交维度,系统分析智能体安全的脆弱性。为此,提出OASIS(Orthogonal Agent Safety Inquiry Suite)——一个分层结构的基准测试体系,包含细粒度标注和高保真仿真沙箱。研究发现:随着意图逐渐隐蔽,安全对齐显著且可预测地下降;同时出现“复杂度悖论”——复杂任务上表现看似更安全,实则源于智能体能力不足。通过开源OASIS及其仿真环境,为探索并强化这些被忽视维度下的智能体安全提供可靠基础。

原文摘要 · Abstract (English)

Current safety evaluations for LLM-driven agents primarily focus on atomic harms, failing to address sophisticated threats where malicious intent is concealed or diluted within complex tasks. We address this gap with a two-dimensional analysis of agent safety brittleness under the orthogonal pressures of intent concealment and task complexity. To enable this, we introduce OASIS (Orthogonal Agent Safety Inquiry Suite), a hierarchical benchmark with fine-grained annotations and a high-fidelity simulation sandbox. Our findings reveal two critical phenomena: safety alignment degrades sharply and predictably as intent becomes obscured, and a "Complexity Paradox" emerges, where agents seem safer on harder tasks only due to capability limitations. By releasing OASIS and its simulation environment, we provide a principled foundation for probing and strengthening agent safety in these overlooked dimensions.

智能体安全意图隐蔽评测基准复杂度悖论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。