让AI自动决定思考长短,又快又准用工具解决问题。
AutoTool: Automatic Scaling of Tool-Use Capabilities in RL via Decoupled Entropy Constraints
- 先用监督学习分清问题难易,再用强化学习自动选择推理长度。
- 在三个基准上准确率提升9.8%,计算开销降低约81%。
- 适合需要高效推理的智能体应用,如复杂任务规划。
工具使用是智能体的关键能力,现有基于强化学习(RL)的方法在扩展推理过程时面临两大挑战:(a) 直接强化学习难以充分扩展思维长度以解决复杂问题;(b) 扩展后的模型对简单问题过度思考,导致大量令牌浪费。为此,我们提出一种新训练范式:首先通过预热监督微调帮助模型区分简单与复杂问题,随后引入强化学习使模型能自动确定合适的推理路径。为实现自动思维长度缩放,我们发现基于熵的优化目标可在保持模型多样性的同时有效释放其扩展潜力。基于此,我们提出一种基于熵的长短推理融合强化学习策略。在三个基准上的实验表明,该模型成功实现了高效的工具使用自动缩放,在准确率提升9.8%的同时,计算开销减少约81%。
原文摘要 · Abstract (English)
Tool use represents a critical capability for AI agents, with recent advances focusing on leveraging reinforcement learning (RL) to scale up the explicit reasoning process to achieve better performance. However, there are some key challenges for tool use in current RL-based scaling approaches: (a) direct RL training often struggles to scale up thinking length sufficiently to solve complex problems, and (b) scaled-up models tend to overthink simpler problems, resulting in substantial token inefficiency. To address these challenges, we propose a novel training paradigm that first employs warm-up supervised fine-tuning to help models distinguish between simple and complex problems, followed by RL that enable models to automatically determine appropriate reasoning trajectories. Furthermore, to tackle the issue of automatic thinking-length scaling, we discover that entropy-based optimization objectives effectively maintain model diversity while successfully unlocking the model's scaling capabilities. Based on this insight, we introduce an entropy-based long-short reasoning fusion RL strategy. Our experiments on three benchmarks demonstrate that model successfully achieves auto-scaling for efficient tool use, achieving significant 9.8\% accuracy improvements while reducing computational overhead by \textasciitilde81\%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。