通过梯度谱分析揭示高质量数据如何影响大模型微调
How Instruction and Reasoning Data shape Post-Training: Data Quality through the Lens of Layer-wise Gradients
- 用梯度奇异值分解分析不同数据质量对微调的影响
- 高质量数据对应更低的核范数和更高的有效秩
- 适合关注数据质量与训练稳定性的研究人员
随着大语言模型后训练从指令跟随向复杂推理任务演进,不同数据对微调动态的影响仍不清楚。本文通过对指令和推理数据引发的层间梯度进行谱分析,发现广泛使用的数据评估指标(如IFD、InsTag、Difficulty、Reward)可由梯度SVD计算出的谱特性统一解释。高质量数据通常具有更低的核范数和更高的有效秩。值得注意的是,有效秩在捕捉细微质量差异方面比核范数更具鲁棒性和分辨力;例如,推理数据的有效秩显著高于指令数据,表明复杂任务带来更丰富的梯度结构。实验还显示,同一家族模型无论规模大小,均呈现相似的梯度模式,而不同家族则显著分化。该研究为指令与推理数据质量的影响提供了统一视角,揭示了数据质量与训练稳定性之间的相互作用,为后训练阶段的数据探索策略提供新思路。
原文摘要 · Abstract (English)
As the post-training of large language models (LLMs) advances from instruction-following to complex reasoning tasks, understanding how different data affect finetuning dynamics remains largely unexplored. In this paper, we present a spectral analysis of layer-wise gradients induced by low/high-quality instruction and reasoning data for LLM post-training. Our analysis reveals that widely-studied metrics for data evaluation, e.g., IFD, InsTag, Difficulty, and Reward, can be explained and unified by spectral properties computed from gradients' singular value decomposition (SVD). Specifically, higher-quality data are usually associated with lower nuclear norms and higher effective ranks. Notably, effective rank exhibits better robustness and resolution than nuclear norm in capturing subtle quality differences. For example, reasoning data achieves substantially higher effective ranks than instruction data, implying richer gradient structures on more complex tasks. Our experiments also highlight that models within the same family share similar gradient patterns regardless of their sizes, whereas different model families diverge significantly. Providing a unified view on the effects of data quality across instruction and reasoning data, this work illuminates the interplay between data quality and training stability, shedding novel insights into developing better data exploration strategies for post-training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。