arXiv:2509.14718cs.LGcs.CL2025-09被引 4

针对工具学习中冗余样本多的问题,提出双动态采样框架提升强化学习效率。

ToolSample: Dual Dynamic Sampling Methods with Curriculum Learning for RL-based Tool Learning

  • 基于奖励统计与任务难度动态筛选样本和子任务
  • 在BFCLv3上比基线提升3.29%性能
  • 适合需要高效训练多任务工具学习模型的研究者

尽管强化学习在基于大模型的工具学习中应用日益广泛,但训练效率常因大量简单样本导致学习价值递减而受阻。现有动态采样方法难以适配工具学习固有的多任务结构与细粒度奖励机制。本文提出面向课程学习的动态采样框架(DSCL),专为工具学习特点设计:多相互依赖的子任务与多值奖励函数。该框架包含两个核心组件:基于奖励的动态采样,利用多维奖励统计量(均值与方差)优先选择高价值数据;基于任务的动态课程学习,自适应聚焦于掌握程度较低的子任务。通过大量实验验证,DSCL显著提升训练效率与模型性能,在BFCLv3基准上实现3.29%的性能提升。该方法有效利用工具学习中的复杂奖励信号与子任务动态,达成更优结果。

原文摘要 · Abstract (English)

While reinforcement learning (RL) is increasingly used for LLM-based tool learning, its efficiency is often hampered by an overabundance of simple samples that provide diminishing learning value as training progresses. Existing dynamic sampling techniques are ill-suited for the multi-task structure and fine-grained reward mechanisms inherent to tool learning. This paper introduces Dynamic Sampling with Curriculum Learning (DSCL), a framework specifically designed to address this challenge by targeting the unique characteristics of tool learning: its multiple interdependent sub-tasks and multi-valued reward functions. DSCL features two core components: Reward-Based Dynamic Sampling, which uses multi-dimensional reward statistics (mean and variance) to prioritize valuable data, and Task-Based Dynamic Curriculum Learning, which adaptively focuses training on less-mastered sub-tasks. Through extensive experiments, we demonstrate that DSCL significantly improves training efficiency and model performance over strong baselines, achieving a 3.29\% improvement on the BFCLv3 benchmark. Our method provides a tailored solution that effectively leverages the complex reward signals and sub-task dynamics within tool learning to achieve superior results.

强化学习工具学习动态采样课程学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。