arXiv:2505.15353cs.CL2025-05被引 2

为语言模型的KL散度建立跨设置统一量纲,揭示训练行为早熟现象。

Establishing a Scale for Kullback-Leibler Divergence in Language Models Across Various Settings

  • 用对数似然向量构建统一比较空间,支持多阶段、多配置模型对比。
  • 发现训练中对数似然空间的KL变化远小于权重空间,呈现亚扩散轨迹。
  • 适用于模型分析、训练监控与跨规模/量化版本的可比性研究。

对数似然向量为语言模型作为概率分布的比较提供统一空间,支持在异构设置下的统一比较。我们将该框架扩展至训练检查点与中间层,建立了预训练、模型规模、随机种子、量化、微调及层间的统一KL散度量纲。对Pythia预训练轨迹的分析表明,对数似然空间中的变化(由KL散度缩放行为衡量)远小于权重空间,导致亚扩散学习轨迹,尽管权重持续漂移,语言模型行为仍提前稳定。

原文摘要 · Abstract (English)

Log-likelihood vectors define a common space for comparing language models as probability distributions, enabling unified comparisons across heterogeneous settings. We extend this framework to training checkpoints and intermediate layers, and establish a consistent scale for KL divergence across pretraining, model size, random seeds, quantization, fine-tuning, and layers. Analysis of Pythia pretraining trajectories further shows that changes in log-likelihood space, as measured by the scaling behavior of KL divergence, are much smaller than in weight space, resulting in subdiffusive learning trajectories and early stabilization of language-model behavior despite weight drift.

KL散度模型训练语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。