arXiv:2606.02965cs.AI2026-06中稿 · AIES 2026, this is…被引 1

提出智能体主动放弃任务的新框架,提升安全性和可控性。

Designing for Doubt: The Case for Informed Abstention in Autonomous Agents

  • 将拒绝执行视为有意识的能力,而非失败
  • 运行时拦截率达87.5%-91%,授权场景可用性达75%-92%
  • 适合关注安全部署的AI系统设计者

随着大语言模型获得工具调用能力并被部署为可自主编辑记录、执行交易和修改基础设施的智能体,当前仍仅以任务完成率作为评价标准,这构成系统性设计缺陷。基准评分、产品指标和默认部署配置均鼓励智能体在缺乏必要输入、证据或授权时仍继续操作,称为‘合规偏差’。本文提出三个贡献:首先,揭示合规偏差如何嵌入现有评估体系——主流基准要么惩罚暂停,要么不衡量暂停是否合理;其次,引入‘知情放弃框架’,将暂停重构为预条件感知的结构化能力:阻断下一步工具调用,明确缺失要素,并引导具体恢复动作;第三,主张运行时强制、校准的防护机制与可审计日志应成为智能体系统设计的标准属性。我们在144个场景和7个模型家族中进行了初步评估,结果显示运行时强制可实现87.5%-91%危险操作拦截,授权场景可用性为75%-92%;不同模型家族中合规偏差呈现两种相反结构;安全与可用性的权衡是可调节的,而非固定不变。

原文摘要 · Abstract (English)

As large language models gain tool access and are deployed as autonomous agents capable of editing records, executing transactions, and modifying infrastructure, we still evaluate them based on the sole metric of task completion. We argue that this evaluation focus constitutes a systematic design failure. Benchmark scoring, product metrics, and default deployment configurations all reward agents for proceeding even when they lack the inputs, evidence, or authorization required to do so safely. We call this phenomenon compliance bias. This paper makes three contributions. First, we show how compliance bias is embedded in the evaluation regimes that currently shape agent development: prominent benchmarks either penalize agents for pausing or fail to measure whether pausing was appropriate. Second, we introduce the Informed Abstention Framework, which reconceptualizes abstention not as a failure mode but as a structured capability: a precondition-aware pause that blocks the next tool call, names what is missing, and routes to a concrete recovery action. Third, we specify what informed abstention requires in deployment, arguing that runtime enforcement, calibrated guard mechanisms, and auditable trace generation should become standard properties of agentic system design rather than optional additions. We perform a preliminary evaluation of our approach across 144 scenarios and seven model families. Our results show that runtime enforcement achieves 87.5-91% hazardous-action blocking and 75-92% usability on authorized scenarios, that compliance bias takes two structurally opposite forms across model families, and that the safety-usability tradeoff is tunable rather than fixed.

智能体安全合规偏差运行时防护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。