arXiv:2508.12116cs.LGcs.AI2025-08ACL被引 4

动态优化指令微调数据集混合比例,提升模型性能。

DynamixSFT: Dynamic Mixture Optimization of Instruction Tuning Collections

  • 将数据混合优化建模为多臂赌博机问题,动态调整采样概率。
  • 在10个基准上显著提升Tulu-2/3模型表现,计算开销极低。
  • 适合需要高效训练且数据源多的AI研发团队使用。

随着大量指令微调数据集不断涌现,动态平衡与优化其混合比例成为关键挑战。为此,我们提出DynamixSFT,一种动态且自动化的指令微调数据集混合优化方法。将问题建模为多臂赌博机,引入先验加权的Boltzmann探索机制,使更新后的采样分布软性锚定于原始数据集比例,从而保持集合固有的多样性与覆盖范围。采样概率通过轻量级一步前瞻奖励进行更新,反映各数据集对当前模型性能提升的贡献度。实验表明,DynamixSFT在10个基准上有效优化了Tulu-2和Tulu-3数据混合集,相比朴素采样仅引入极小计算开销。此外,我们提供了全面分析与可视化,深入揭示该方法的自适应动态机制。

原文摘要 · Abstract (English)

As numerous instruction-tuning datasets continue to emerge, dynamically balancing and optimizing their mixtures has become a critical challenge. To address this, we propose DynamixSFT, a dynamic and automated method for instruction-tuning dataset mixture optimization. We formulate the problem as a multi-armed bandit setup and introduce a Prior-scaled Boltzmann Exploration that softly anchors the updated sampling distribution to the original dataset proportions, thereby preserving the inherent diversity and coverage of the collection. Sampling probabilities are updated using a lightweight 1-Step Look-ahead Reward, reflecting how much the dataset contributes to improving the model's performance at its current state. We demonstrate that DynamixSFT effectively optimizes the Tulu-2-mixture and Tulu-3-mixture collections across 10 benchmarks, while introducing minimal computational overhead over naive sampling. Furthermore, we provide a comprehensive analysis and visualizations to offer deeper insights into the adaptive dynamics of our method.

指令微调数据混合强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。