arXiv:2505.06150cs.CLcs.AI2025-05被引 1

提出新规律:固定算力下,数据组成影响微调效率

A Scaling Law for Token Efficiency in LLM Fine-Tuning Under Fixed Compute Budgets

  • 用数据量和平均长度共同衡量数据组成,替代单一总词元数
  • 实验表明不同数据组合下,相同算力下的性能差异显著
  • 适合资源受限场景下优化大模型微调策略的研究者

我们提出了在固定算力预算下微调大语言模型的缩放定律,明确考虑了数据构成的影响。传统方法仅以总词元数衡量训练数据,但样本数量及其平均词元长度——我们称之为数据集容量——对模型性能具有决定性作用。该公式遵循既定调优流程进行校准。在BRICC数据集及MMLU数据集子集上,通过多种采样策略评估发现,数据构成显著影响词元效率。这些结果推动了在资源受限环境下实用的大语言模型微调缩放定律的改进。

原文摘要 · Abstract (English)

We introduce a scaling law for fine-tuning large language models (LLMs) under fixed compute budgets that explicitly accounts for data composition. Conventional approaches measure training data solely by total tokens, yet the number of examples and their average token length -- what we term \emph{dataset volume} -- play a decisive role in model performance. Our formulation is tuned following established procedures. Experiments on the BRICC dataset \cite{salavati2024reducing} and subsets of the MMLU dataset \cite{hendrycks2021measuringmassivemultitasklanguage}, evaluated under multiple subsampling strategies, reveal that data composition significantly affects token efficiency. These results motivate refined scaling laws for practical LLM fine-tuning in resource-constrained settings.

大模型微调缩放定律数据效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。