为工具调用大模型设计专用奖励模型,提升执行效果与鲁棒性。
ToolRM: Outcome Reward Models for Tool-Calling Large Language Models
- 基于开源大模型合成数据,训练专用于工具调用的奖励模型。
- 在多种场景下性能优于通用模型,最佳采样提升达25%。
- 适合需要可靠工具执行的强化学习和数据筛选任务。
随着大语言模型越来越多地与外部工具交互,针对工具使用的奖励建模已成为关键但研究不足的领域。现有奖励模型主要基于自然语言输出训练,难以评估工具推理与执行效果。为此,我们提出首个系统性评估工具调用场景中奖励模型的基准测试FC-RewardBench。分析表明,当前模型常忽略有效工具使用的关键信号,凸显了领域专用建模的必要性。我们通过利用许可宽松的开源大模型生成数据,提出一种训练框架,构建了ToolRM——一套参数规模从1.7B到14B的工具调用奖励模型。在多种设置下,这些模型均显著优于通用基线,尤其在使用Best-of-N采样时性能提升最高达25%,同时增强对输入噪声的鲁棒性,支持高效数据过滤,并可直接用于策略模型的强化学习训练。
原文摘要 · Abstract (English)
As large language models (LLMs) increasingly interact with external tools, reward modeling for tool use has emerged as a critical yet underexplored area of research. Existing reward models, trained primarily on natural language outputs, struggle to evaluate tool-based reasoning and execution. To quantify this gap, we introduce FC-RewardBench, the first benchmark to systematically evaluate reward models in tool-calling scenarios. Our analysis shows that current reward models frequently miss key signals of effective tool use, highlighting the need for domain-specific modeling. We address this by proposing a training framework for outcome reward models using data synthesized from permissively licensed, open-weight LLMs. We introduce ToolRM - a suite of reward models for tool-use ranging from 1.7B to 14B parameters. Across diverse settings, these models consistently outperform general-purpose baselines. Notably, they achieve up to a 25% improvement with Best-of-N sampling, while also improving robustness to input noise, enabling effective data filtering, and supporting RL-training of policy models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。