arXiv:2605.30524cs.LG2026-05被引 1

发现大模型多阶段微调会压缩表征,导致学习能力下降。

Representation Collapse in Sequential Post-Training of Large Language Models

论文配图:Representation Collapse in Sequential Post-Training of Large Language Models
图 1 · 摘自论文原文
  • 通过多维度测量分析表征演化过程。
  • 多阶段微调使特征空间趋于低秩同质化,影响泛化与适应性。
  • 提出轻量干预方法保留未来可学习性,适合持续训练场景。

大语言模型现常通过多个后训练阶段进行适配,而非单一指令微调。本文研究这种序列式后训练是否会逐步将内部表征压缩至低秩、各向异性和同质的特征空间。我们定义了一套针对隐藏状态、输出逻辑值、标记轨迹及LoRA更新的测量体系,并在受控阶段顺序下分析了监督微调、偏好优化、安全/拒答调优、数学与代码专精、长链思维推理调优。核心假设是:过度表征集中不仅是几何现象,还会预测后续适应性减弱、域外泛化能力下降和校准性能变差。我们进一步评估了轻量级干预手段,包括跨领域重放、特征刷新、表征多样性正则化及LoRA更新去相关性,以在不牺牲后训练行为收益的前提下,维持未来可学习性。

原文摘要 · Abstract (English)

Large language models are now adapted through chains of post-training stages rather than through a single instruction-tuning pass. This paper studies whether such sequential post-training gradually compresses internal representations into low-rank, anisotropic, and homogeneous feature spaces. We define a measurement suite for hidden states, logits, token trajectories, and LoRA updates, and we use it to analyze supervised fine-tuning, preference optimization, safety/refusal tuning, math and code specialization, and long chain-of-thought tuning under controlled stage orderings. The central hypothesis is that excessive representation concentration is not merely a geometric curiosity: it predicts reduced plasticity during later adaptation, weaker out-of-domain generalization, and poorer calibration. We further evaluate lightweight interventions, including mixed-domain replay, feature refresh, representation diversity regularization, and LoRA update decorrelation, as ways to preserve future learnability without giving up the behavioral gains of post-training.

表征压缩大模型训练持续学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。