arXiv:2510.18383cs.CLcs.AI2025-10被引 1

让小模型学会灵活用工具,比死记硬背更有效。

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation

  • 用动态奖励引导小模型模仿大模型的工具使用过程
  • 在验证环境中,泛化能力显著优于传统方法
  • 适合需要灵活适应新任务的小模型应用

将大语言模型(LLM)的工具使用能力蒸馏到小语言模型(SLM)中,对实际应用至关重要。主流的监督微调(SFT)是离策略蒸馏方法,因僵化对齐静态教师轨迹,导致域外(OOD)泛化性能差。尽管强化学习(RL)是替代方案,但小模型能力受限:稀疏结果奖励难以提供有效指导,而严格轨迹匹配又施加过强约束。为弥合这一能力差距,我们提出 MENTOR,一种基于策略的蒸馏框架,引入灵活且过程感知的奖励结构。不同于强制复制,MENTOR利用教师参考来引导工具使用行为,在行为对齐与下游性能间取得平衡。在可控可执行工具基准上的大量实验表明,相较于 SFT 和严格 RL 基线,MENTOR 显著提升域外工具使用性能。研究发现,在可验证的工具使用环境中,灵活对齐比严格轨迹复制更有效,有助于构建适应性强的小模型。

原文摘要 · Abstract (English)

Distilling the tool-use capabilities of large language models (LLMs) into small language models (SLMs) is essential for their practical application. The predominant approach, supervised fine-tuning (SFT), is an off-policy distillation method that suffers from poor out-of-domain (OOD) generalization because it rigidly aligns with static teacher trajectories. While reinforcement learning (RL) offers an alternative, the capacity limitations of SLMs pose a severe dilemma: sparse outcome rewards provide insufficient guidance, whereas strict trajectory matching imposes overly restrictive constraints. To bridge this capacity-driven gap, we propose MENTOR, an on-policy distillation framework that introduces a flexible yet process-aware reward structure. Instead of enforcing rigid replication, MENTOR uses the teacher's reference to guide tool-use behavior, balancing behavioral alignment with downstream performance. Extensive experiments on controlled executable-tool benchmarks demonstrate that MENTOR improves OOD tool-use performance compared to SFT and strict RL baselines. Our findings suggest that within verifiable tool-use environments, flexible tool-use alignment offers a more effective approach than strict trajectory replication for developing adaptable small models.

模型蒸馏强化学习工具使用小模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。