arXiv:2503.23383cs.CL2025-03被引 122

让大模型通过强化学习自主使用工具解题,效果远超传统方法。

ToRL: Scaling Tool-Integrated RL

  • 用强化学习训练模型自主探索工具使用策略。
  • 70亿参数模型在AIME24上达43.3%准确率,领先17%。
  • 能自发调整策略,动态切换计算与分析思维。

我们提出ToRL(工具集成强化学习)框架,训练大语言模型通过强化学习自主使用计算工具。与监督微调不同,ToRL使模型能够探索并发现最优工具使用策略。实验显示,Qwen2.5-Math系列模型在AIME 24测试中表现显著提升:ToRL-7B达到43.3%准确率,较无工具集成的强化学习高出14%,优于当前最佳工具集成推理(TIR)模型17%。进一步分析揭示了涌现行为,包括策略性调用工具、自我抑制无效代码、以及在计算与分析推理间动态适应,这些均源于奖励驱动的学习过程。

原文摘要 · Abstract (English)

We introduce ToRL (Tool-Integrated Reinforcement Learning), a framework for training large language models (LLMs) to autonomously use computational tools via reinforcement learning. Unlike supervised fine-tuning, ToRL allows models to explore and discover optimal strategies for tool use. Experiments with Qwen2.5-Math models show significant improvements: ToRL-7B reaches 43.3\% accuracy on AIME~24, surpassing reinforcement learning without tool integration by 14\% and the best existing Tool-Integrated Reasoning (TIR) model by 17\%. Further analysis reveals emergent behaviors such as strategic tool invocation, self-regulation of ineffective code, and dynamic adaptation between computational and analytical reasoning, all arising purely through reward-driven learning.

强化学习工具使用大模型数学推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。