用熵减提升大模型工具使用效率,减少无效调用。
Rethinking the Role of Entropy in Optimizing Tool-Use Behaviors for Large Language Model Agents
- 以熵减作为监督信号,引导高质量工具调用。
- 减少72.07%工具调用次数,性能提升22.27%。
- 适合需要高效推理的智能体应用场景。
基于大语言模型的工具使用智能体在数学推理和多跳问答等任务中表现优异。然而,在长轨迹任务中,智能体常触发过多且低质量的工具调用,增加延迟并降低推理效率,难以有效管理工具使用行为。本文通过熵基预实验发现,熵降低与高质量工具调用存在强正相关。基于此,提出将熵减作为监督信号,设计两种奖励策略:稀疏结果奖励提供粗粒度轨迹级指导以提升效率,密集过程奖励则提供细粒度监督以增强性能。跨多个领域的实验表明,两种策略均有效:前者使工具调用次数相比基线平均值减少72.07%,后者性能提升22.27%。这些结果表明熵减是优化工具使用行为的关键机制,使智能体在真实应用中更具适应性。
原文摘要 · Abstract (English)
Tool-using agents based on Large Language Models (LLMs) excel in tasks such as mathematical reasoning and multi-hop question answering. However, in long trajectories, agents often trigger excessive and low-quality tool calls, increasing latency and degrading inference performance, making managing tool-use behavior challenging. In this work, we conduct entropy-based pilot experiments and observe a strong positive correlation between entropy reduction and high-quality tool calls. Building on this finding, we propose using entropy reduction as a supervisory signal and design two reward strategies to address the differing needs of optimizing tool-use behavior. Sparse outcome rewards provide coarse, trajectory-level guidance to improve efficiency, while dense process rewards offer fine-grained supervision to enhance performance. Experiments across diverse domains show that both reward designs improve tool-use behavior: the former reduces tool calls by 72.07% compared to the average of baselines, while the latter improves performance by 22.27%. These results position entropy reduction as a key mechanism for enhancing tool-use behavior, enabling agents to be more adaptive in real-world applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。