arXiv:2510.26167cs.AIcs.CL2025-10ACL被引 3

为工具调用设计轻量级奖励模型,提升智能体任务准确率。

ToolRM: Towards Agentic Tool-Use Reward Modeling

  • 基于规则评分与多维采样构建高质量偏好数据集
  • 在30K条数据上训练,工具调用准确率提升17.94%
  • 支持推理时扩展,减少66%以上输出token

奖励模型(RMs)在对齐大语言模型(LLMs)与人类偏好方面起着关键作用。然而,在工具学习领域,缺乏专为函数调用任务设计的奖励模型,限制了更强大智能体AI的发展。我们提出ToolRM,一类针对通用工具使用场景的轻量级奖励模型。为构建这些模型,我们提出一种新流程,利用基于规则的评分和多维采样生成高质量成对偏好数据。该流程生成了ToolPref-Pairwise-30K数据集,具有多样性、平衡性和挑战性,支持生成式与判别式奖励建模。我们还引入TRBench$_{BFCL}$,基于代理评估套件BFCL构建的基准测试,用于评估工具调用任务中的奖励模型。在所构建数据上训练的Qwen3-4B/8B系列模型,在成对奖励判断中准确率最高提升17.94%,显著优于前沿的LLMs和奖励模型。除了训练目标外,生成式ToolRM还能泛化至更广泛的批判任务,包括Best-of-N采样与自纠错。在ACEBench上的实验表明其高效有效,实现推理时扩展的同时,输出token使用减少超过66%。其对下游强化学习训练的支持进一步验证了其实际应用价值。我们已开源相关数据以促进未来研究。

原文摘要 · Abstract (English)

Reward models (RMs) play a critical role in aligning large language models (LLMs) with human preferences. Yet in the domain of tool learning, the lack of RMs specifically designed for function-calling tasks has limited progress toward more capable agentic AI. We introduce ToolRM, a family of lightweight reward models tailored for general tool-use scenarios. To build these models, we propose a novel pipeline that constructs high-quality pairwise preference data using rule-based scoring and multidimensional sampling. This yields ToolPref-Pairwise-30K, a diverse, balanced, and challenging preference dataset that supports both generative and discriminative reward modeling. We also introduce TRBench$_{BFCL}$, a benchmark built on the agent evaluation suite BFCL to evaluate RMs on tool calling tasks. Trained on our constructed data, models from the Qwen3-4B/8B series achieve up to 17.94% higher accuracy, substantially outperforming frontier LLMs and RMs in pairwise reward judgments. Beyond training objectives, generative ToolRM generalizes to broader critique tasks, including Best-of-N sampling and self-correction. Experiments on ACEBench highlight its effectiveness and efficiency, enabling inference-time scaling while reducing output token usage by over 66%. Its support for downstream RL training further validates its practical utility. We release data to facilitate future research.

奖励模型工具调用智能体Qwen

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。