arXiv:2510.16492cs.CL2025-10被引 13

让大模型主动放弃不确定任务,显著提升安全性。

Check Yourself Before You Wreck Yourself: Selectively Quitting Improves LLM Agent Safety

  • 通过显式指令让模型在不自信时退出任务
  • 安全得分平均提升0.39,帮助性仅下降0.03
  • 适合高风险场景的自主智能体使用

随着大语言模型代理在复杂环境中执行具有现实后果的任务,其安全性变得至关重要。尽管不确定性量化在单轮任务中已有深入研究,但在涉及真实工具访问的多轮代理场景中,不确定性与模糊性会累积,带来远超传统文本生成失败的严重甚至灾难性风险。我们提出将‘退出’作为一种简单而有效的行为机制,使大模型代理在缺乏信心时能识别并主动退出。基于ToolEmu框架,我们在12个顶尖大模型上系统评估了退出行为。结果表明,给予显式退出指令后,所有模型的安全性平均提升0.39(0-3分制),专有模型提升达0.64;同时帮助性仅轻微下降0.03。分析显示,仅添加显式退出指令即可立即部署于现有代理系统,可作为高风险应用中自主代理的第一道安全防线。

原文摘要 · Abstract (English)

As Large Language Model (LLM) agents increasingly operate in complex environments with real-world consequences, their safety becomes critical. While uncertainty quantification is well-studied for single-turn tasks, multi-turn agentic scenarios with real-world tool access present unique challenges where uncertainties and ambiguities compound, leading to severe or catastrophic risks beyond traditional text generation failures. We propose using "quitting" as a simple yet effective behavioral mechanism for LLM agents to recognize and withdraw from situations where they lack confidence. Leveraging the ToolEmu framework, we conduct a systematic evaluation of quitting behavior across 12 state-of-the-art LLMs. Our results demonstrate a highly favorable safety-helpfulness trade-off: agents prompted to quit with explicit instructions improve safety by an average of +0.39 on a 0-3 scale across all models (+0.64 for proprietary models), while maintaining a negligible average decrease of -0.03 in helpfulness. Our analysis demonstrates that simply adding explicit quit instructions proves to be a highly effective safety mechanism that can immediately be deployed in existing agent systems, and establishes quitting as an effective first-line defense mechanism for autonomous agents in high-stakes applications.

大模型安全智能体退出机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。