arXiv:2509.23024cs.LGcs.AI2025-09NeurIPS被引 33

通过谱分析揭示大模型从预训练到微调的表征几何演化规律。

Tracing the Representation Geometry of Language Models from Pretraining to Post-training

  • 用有效秩和特征谱衰减速率量化表征几何变化。
  • 发现预训练中存在三阶段几何演化:坍缩→扩张→压缩,对应能力涌现。
  • 微调阶段不同方法会引发不同几何动态,影响模型性能与鲁棒性。

标准训练指标如损失值无法解释大语言模型复杂能力的涌现。本文采用谱方法研究自回归预训练及后训练过程中学习表征的几何特性,通过有效秩(RankMe)和特征谱衰减速率(α-ReQ)进行测量。基于OLMo(1B-7B)和Pythia(160M-12B)模型,我们发现预训练中存在一致的非单调三阶段几何演化:初始‘热身’阶段出现快速表征坍缩;随后进入‘熵寻求’阶段,流形维度显著扩展,与最高阶词组记忆峰值同步;最后是‘压缩寻求’阶段,实施各向异性整合,选择性保留主导特征方向方差而压缩其他方向,此转变与下游任务性能显著提升相关。该现象源于交叉熵优化在偏斜词频分布和表示瓶颈(d ≪ |V|)下的基本相互作用。后训练进一步改变几何结构:SFT与DPO驱动‘熵寻求’动态以融合特定指令或偏好数据,提升分布内性能但降低分布外鲁棒性;而RLVR则引发‘压缩寻求’,增强奖励对齐但减少生成多样性。

原文摘要 · Abstract (English)

Standard training metrics like loss fail to explain the emergence of complex capabilities in large language models. We take a spectral approach to investigate the geometry of learned representations across pretraining and post-training, measuring effective rank (RankMe) and eigenspectrum decay ($α$-ReQ). With OLMo (1B-7B) and Pythia (160M-12B) models, we uncover a consistent non-monotonic sequence of three geometric phases during autoregressive pretraining. The initial "warmup" phase exhibits rapid representational collapse. This is followed by an "entropy-seeking" phase, where the manifold's dimensionality expands substantially, coinciding with peak n-gram memorization. Subsequently, a "compression-seeking" phase imposes anisotropic consolidation, selectively preserving variance along dominant eigendirections while contracting others, a transition marked with significant improvement in downstream task performance. We show these phases can emerge from a fundamental interplay of cross-entropy optimization under skewed token frequencies and representational bottlenecks ($d \ll |V|$). Post-training further transforms geometry: SFT and DPO drive "entropy-seeking" dynamics to integrate specific instructional or preferential data, improving in-distribution performance while degrading out-of-distribution robustness. Conversely, RLVR induces "compression-seeking", enhancing reward alignment but reducing generation diversity.

表征几何预训练微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。