arXiv:2601.07264cs.CL2026-01ACL被引 10

发现工具类型导致智能体自信偏差,提出强化学习优化方法提升可信度。

The Confidence Dichotomy: Analyzing and Mitigating Miscalibration in Tool-Use Agents

  • 按工具类型区分校准策略:检索类工具易过自信,验证类工具可抑制偏差。
  • 用强化学习联合优化任务准确率与校准性,在多场景下显著提升可靠性。
  • 适合关注AI可信部署、安全决策的开发者与研究者参考。

基于大语言模型的自主智能体正快速发展以处理多轮任务,但确保其可信性仍是关键挑战。校准能力——即智能体表达的自信程度与其实际表现的一致性——是可信性的核心。尽管静态模型的校准已较为成熟,但工具集成的智能体工作流中的校准动态仍不清晰。本文系统研究了工具使用智能体中的言语化校准,揭示出由工具类型驱动的根本性信心二分现象:证据类工具(如网络搜索)因检索信息固有的噪声,系统性引发严重过自信;而验证类工具(如代码解释器)可通过确定性反馈实现推理锚定,缓解校准偏差。为在各类工具间稳健提升校准性,我们提出一种强化学习微调框架,联合优化任务准确率与校准性,并构建了涵盖多种奖励设计的综合评估基准。实验表明,训练后的智能体不仅校准性能更优,且在从局部训练环境泛化至嘈杂网络环境,以及跨数学推理等不同领域时均表现出强鲁棒性。结果强调了针对工具类型的专用校准策略的必要性。更广泛而言,本工作为构建能可靠传达不确定性的自我意识智能体奠定了基础,适用于高风险真实场景部署。

原文摘要 · Abstract (English)

Autonomous agents based on large language models (LLMs) are rapidly evolving to handle multi-turn tasks, but ensuring their trustworthiness remains a critical challenge. A fundamental pillar of this trustworthiness is calibration, which refers to an agent's ability to express confidence that reliably reflects its actual performance. While calibration is well-established for static models, its dynamics in tool-integrated agentic workflows remain underexplored. In this work, we systematically investigate verbalized calibration in tool-use agents, revealing a fundamental confidence dichotomy driven by tool type. Specifically, our pilot study identifies that evidence tools (e.g., web search) systematically induce severe overconfidence due to inherent noise in retrieved information, while verification tools (e.g., code interpreters) can ground reasoning through deterministic feedback and mitigate miscalibration. To robustly improve calibration across tool types, we propose a reinforcement learning (RL) fine-tuning framework that jointly optimizes task accuracy and calibration, supported by a holistic benchmark of reward designs. We demonstrate that our trained agents not only achieve superior calibration but also exhibit robust generalization from local training environments to noisy web settings and to distinct domains such as mathematical reasoning. Our results highlight the necessity of domain-specific calibration strategies for tool-use agents. More broadly, this work establishes a foundation for building self-aware agents that can reliably communicate uncertainty in high-stakes, real-world deployments.

智能体校准大模型可信AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。