数据难易度影响大模型微调效果,需根据数据量动态选择。
Data Difficulty and the Generalization--Extrapolation Tradeoff in LLM Fine-Tuning

- 基于数据难易度筛选样本,结合数据量调整最优难度。
- 数据量越大,越应选择较难样本以提升性能。
- 适用于需要高效数据筛选的模型微调场景。
监督微调(SFT)中的数据选择会显著影响大语言模型(LLM)的行为。尽管已有研究探讨了基于困惑度、难易度或长度等启发式方法的数据选择,但结果常不一致且依赖上下文。本文从实验和理论两方面系统研究了数据难易度的作用,发现不存在普适最优难易度,其有效性取决于数据集规模。在固定数据预算下,存在一个最优数据难易度,且该最优值随数据预算增加向更难数据偏移。通过受控合成实验揭示,这一现象源于分布内泛化差距与外推差距的相互作用。进一步利用PAC-Bayesian泛化界进行理论分析,验证了该机制。研究阐明了数据量与难易度如何共同影响SFT中泛化与外推的权衡,为特定条件下的难度导向数据选择提供指导。
原文摘要 · Abstract (English)
Data selection during supervised fine-tuning (SFT) can critically change the behavior of large language models (LLMs). Although existing work has studied the effect of selecting data based on heuristics such as perplexity, difficulty, or length, the reported findings are often inconsistent or context-dependent. In this work, we systematically study the role of data difficulty in fine-tuning from both empirical and theoretical perspectives, and find that there is no universally optimal difficulty level; rather, its effectiveness depends on the dataset size. We show that for a fixed data budget, there exists an optimal data difficulty for SFT, and that this optimal difficulty shifts toward harder data as the data budget increases. To explain this phenomenon, we conduct controlled synthetic experiments that reveal a simple underlying mechanism: the interplay between the (in-distribution) generalization gap and the extrapolation gap. We further support this mechanism through a theoretical analysis using PAC-Bayesian generalization bounds. Overall, our results clarify how data size and difficulty jointly affect the trade-off between generalization and extrapolation in SFT, providing guidance for difficulty-based data selection under certain model and data conditions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。