arXiv:2507.11371cs.LGcs.MA2025-07

让大模型学会用冷门工具,提升推理多样性而不降准确率。

Step-wise Policy for Rare-tool Knowledge (SPaRK): Offline RL that Drives Diverse Tool Use in LLMs

  • 用双重奖励机制,同时优化答案质量和工具使用多样性。
  • 在14个MMLU-Pro类别上表现媲美基线,工具选择熵值显著更高。
  • 适合想提升模型探索能力、避免工具使用单一化的研究者。

我们提出一种名为SPaRK的新型强化学习框架,旨在让大语言模型在离线强化学习中探索多样化的工具使用模式,突破传统高温度采样的局限。基于近期的分步强化学习进展,我们设计了一种双目标奖励系统,同时优化答案质量与工具多样性,在合成轨迹上使用离线PPO训练了Llama-3.1 8B模型,数据源为MMLU-Pro。该方法采用罕见优先的利用策略:由GPT-4o裁判评估八种不同工具及思维链推理的候选动作,策略倾向于选择使用频率较低但依然可行的工具,以促进系统性探索。实验结果表明,SPaRK在14个MMLU-Pro类别上达到具有竞争力的表现,同时工具选择的熵值显著高于基线和监督微调方法,说明通过显式工具多样性驱动的算法探索可在不牺牲准确性的前提下增强推理能力。

原文摘要 · Abstract (English)

We present Step-wise Policy for Rare-tool Knowledge (SPaRK), a novel reinforcement learning framework that teaches large language models to explore diverse tool usage patterns beyond conventional high-temperature sampling. Building on recent advances in step-wise reinforcement learning, we introduce a dual-objective reward system that simultaneously optimizes for answer quality and tool diversity, training a Llama-3.1 8B model through offline PPO on synthetically generated trajectories from the MMLU-Pro dataset. Our approach uniquely employs a rarity-first exploitation strategy where a GPT-4o judge scores candidate actions across eight distinct tools plus chain-of-thought reasoning, with the policy favoring less-frequently used but still viable tools to encourage systematic exploration. Empirical results demonstrate that SPaRK achieves competitive performance across 14 MMLU-Pro categories while exhibiting significantly higher entropy in tool selection compared to both baseline and supervised fine-tuning approaches, suggesting that algorithmic exploration through explicit tool diversity can enhance reasoning capabilities without sacrificing accuracy.

强化学习工具使用推理增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。