arXiv:2505.01523cs.LGcs.AI2025-05被引 1

用精选数据提升数学领域大模型微调效率,效果接近全量训练

Subset Selection for Fine-Tuning: A Utility-Diversity Balanced Approach for Mathematical Domain Adaptation

  • 结合性能与多样性,筛选最具代表性的数学题
  • 仅用部分数据即达近似全量训练效果,节省大量算力
  • 适合资源有限但需高效适配数学领域的研究者

我们提出一种精细化的预算型子集选择方法,用于高效微调大语言模型(LLM)在数学领域。该方法融合效用与多样性指标,选取最具信息量且覆盖广泛的训练样本。目标是在仅使用部分数据的情况下,实现接近全数据集的性能,同时显著降低计算成本与训练时间。效用度量综合困惑度与思维链(CoT)损失,识别对模型学习最有益的难题;多样性度量则确保各数学子领域得到充分覆盖。我们在 LLaMA-3 8B 和 Phi-3 模型上评估,对比随机采样、基于多样性的采样及现有先进子集选择方法,验证了所提方法的有效性。

原文摘要 · Abstract (English)

We propose a refined approach to efficiently fine-tune large language models (LLMs) on specific domains like the mathematical domain by employing a budgeted subset selection method. Our approach combines utility and diversity metrics to select the most informative and representative training examples. The final goal is to achieve near-full dataset performance with meticulously selected data points from the entire dataset while significantly reducing computational cost and training time and achieving competitive performance as the full dataset. The utility metric incorporates both perplexity and Chain-of-Thought (CoT) loss to identify challenging examples that contribute most to model learning, while the diversity metric ensures broad coverage across mathematical subdomains. We evaluate our method on LLaMA-3 8B and Phi-3 models, comparing against several baseline approaches, including random selection, diversity-based sampling, and existing state-of-the-art subset selection techniques.

大模型微调数学推理子集选择

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。