arXiv:2606.30077cs.LGcs.AI2026-06

用高斯过程动态选优质指令数据,提升训练效率与稳定性。

Online Data Selection for Instruction Tuning via Gaussian Processes

论文配图:Online Data Selection for Instruction Tuning via Gaussian Processes
图 1 · 摘自论文原文
  • 基于高斯过程建模语义空间中的数据价值连续分布
  • 在三个数据集上优于现有最优方法,显著提升性能
  • 适合追求高效高质量指令微调的开发者与研究者

随着大语言模型预训练与微调从数据量转向数据质量,高质量数据选择成为关键课题。现有在线数据选择方法多为‘批处理受限’,仅优化随机批次内的局部效用。为此,我们提出GAIA(基于高斯过程的全局自适应指令微调框架),将数据估值建模为全局估计过程。GAIA利用高斯过程回归,在语义空间中建模连续的效用曲面,并通过自适应策略融合机制动态优先选择高价值样本。将策略后验更新视为经典固定共享赫德框架下的专家追踪问题,继承了动态遗憾保证,刻画了训练中非平稳质量评分下的鲁棒性。在三个数据集上的实证评估表明,GAIA显著优于当前最优基线如\greats,证明其在高效指令微调中具备可扩展性与鲁棒性。

原文摘要 · Abstract (English)

With Large Language Model (LLM) pre-training and fine-tuning shifting its focus from data volume to data quality, quality data selection has emerged as a critical research topic. Existing online data selection methods for LLM training are typically "batch-constrained", limiting optimization to local utility within random batches. To overcome this, we propose GAIA (Global Adaptive Instruction tuning via GAussian processes), a framework that formulates data valuation as a global estimation process. GAIA employs Gaussian Process regression to model continuous utility manifolds across the semantic space, utilizing an adaptive strategy fusion mechanism to dynamically prioritize high-utility samples. By casting the strategy-posterior update as an instance of the classical fixed-share Hedge framework for tracking the best expert, we inherit a dynamic-regret guarantee that characterizes GAIA's robustness under non-stationary quality scores during training. Empirical evaluations on three datasets demonstrate that GAIA significantly outperforms state-of-the-art baselines like \greats, establishing our method as a scalable and robust solution for efficient instruction tuning.

指令微调数据选择高斯过程在线学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。