通过不确定性对齐强化学习,提升大模型工具调用的决策准确性和可靠性。
Exploring Agentic Tool-Calling Decisions via Uncertainty-Aligned Reinforcement Learning

- 在奖励设计中引入不确定性量化,作为保持正确与错误动作分离的排斥力。
- 在多个工具调用基准上显著提升决策质量,同时维持更可靠的不确定性估计。
- 适用于需要高可靠决策的多步交互场景,如智能助手、自动化流程。
基于大语言模型的智能体在工具使用决策中常出现无效调用或幻觉直接回答,导致多步交互中误差累积。现有方法主要依赖推理时修正或基于结果的粗粒度奖励,未充分挖掘决策不确定性的特性。我们发现,面向决策的强化学习会削弱正确与错误动作间的不确定性分离,导致过度自信的错误和弱探索信号。为此,我们提出TRUST,将不确定性量化纳入奖励设计,作为维持不确定性分离的排斥力,并采用轻量级关键转折点标注,统一后训练多轮轨迹。在多个工具使用基准上的实验表明,TRUST持续提升了决策质量和代理性能,同时在优化过程中保持更可靠的不确定性估计。
原文摘要 · Abstract (English)
Large language model (LLM)-based agents often make suboptimal tool-use decisions, including unsupported tool invocation and hallucinated direct responses, which may accumulate errors throughout multi-step interactions. Existing approaches mainly improve these behaviors through inference-time correction or coarse-grained reward signals based on decision outcomes and structured checklists, leaving the uncertainty characteristics of agent decisions underexplored. We observe that decision-oriented reinforcement learning tends to weaken the uncertainty separation between correct and incorrect actions, resulting in overconfident mistakes and weaker exploration signals. Therefore, we propose TRUST, which incorporates uncertainty quantification into reward design as a repulsive force for maintaining uncertainty separation, and labels lightweight key-turn annotations for unified post-training of multi-turn trajectories. Experimental results across diverse tool-use benchmarks show that TRUST consistently enhances both decision quality and agent performance while maintaining more reliable uncertainty estimates during optimization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。