动态优化指令微调数据集混合比例,提升模型性能。
DynamixSFT: Dynamic Mixture Optimization of Instruction Tuning Collections
- 将数据混合优化建模为多臂赌博机问题,动态调整采样概率。
- 在10个基准上显著提升Tulu-2/3模型表现,计算开销极低。
- 适合需要高效训练且数据源多的AI研发团队使用。
随着大量指令微调数据集不断涌现,动态平衡与优化其混合比例成为关键挑战。为此,我们提出DynamixSFT,一种动态且自动化的指令微调数据集混合优化方法。将问题建模为多臂赌博机,引入先验加权的Boltzmann探索机制,使更新后的采样分布软性锚定于原始数据集比例,从而保持集合固有的多样性与覆盖范围。采样概率通过轻量级一步前瞻奖励进行更新,反映各数据集对当前模型性能提升的贡献度。实验表明,DynamixSFT在10个基准上有效优化了Tulu-2和Tulu-3数据混合集,相比朴素采样仅引入极小计算开销。此外,我们提供了全面分析与可视化,深入揭示该方法的自适应动态机制。
原文摘要 · Abstract (English)
As numerous instruction-tuning datasets continue to emerge, dynamically balancing and optimizing their mixtures has become a critical challenge. To address this, we propose DynamixSFT, a dynamic and automated method for instruction-tuning dataset mixture optimization. We formulate the problem as a multi-armed bandit setup and introduce a Prior-scaled Boltzmann Exploration that softly anchors the updated sampling distribution to the original dataset proportions, thereby preserving the inherent diversity and coverage of the collection. Sampling probabilities are updated using a lightweight 1-Step Look-ahead Reward, reflecting how much the dataset contributes to improving the model's performance at its current state. We demonstrate that DynamixSFT effectively optimizes the Tulu-2-mixture and Tulu-3-mixture collections across 10 benchmarks, while introducing minimal computational overhead over naive sampling. Furthermore, we provide a comprehensive analysis and visualizations to offer deeper insights into the adaptive dynamics of our method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。