首次系统研究大模型长上下文持续预训练的动态过程。
Revealing the Learning Dynamics of Long-Context Continual Pre-training
- 构建分层分析框架,从行为、概率到机制多维度追踪训练进展。
- 发现800亿参数模型需超1500亿词元训练才达内在饱和。
- 提出用检索头注意力分数监控训练稳定性,比传统指标更可靠。
现有长上下文持续预训练(LCCP)研究多集中于小规模模型和有限数据(数十亿词元)。我们指出,将此类设置直接应用于工业级模型可能导致适应不足和过早终止。此外,当前评估依赖下游任务(如针在草堆中测试),常无法反映内在收敛状态,导致“虚假饱和”。本文首次使用工业级模型 Hunyuan-A13B(共800亿参数),对2000亿词元训练轨迹进行系统性研究。提出分层分析框架,涵盖行为(监督微调探测)、概率(困惑度)与机制(注意力模式)三个层面。发现:(1) 大规模数据必要性:数十亿词元训练不足以支撑工业级模型的LCCP(如Hunyuan-A13B在超过1500亿词元后才达到饱和);(2) 虚假饱和与内在饱和之别:传统NIAH分数过早显示“饱和”,而基于困惑度的分析揭示持续内在提升,且与下游性能相关性更强;(3) 机制监控提升训练稳定性:检索头注意力分数可作为高效低资源的训练监控信号,其变化与微调结果高度相关。本工作为工业级大模型的LCCP提供了全面的监控框架、评估体系与机制解释。
原文摘要 · Abstract (English)
Existing studies on Long-Context Continual Pre-training (LCCP) mainly focus on small-scale models and limited data regimes (tens of billions of tokens). We argue that directly migrating these small-scale settings to industrial-grade models risks insufficient adaptation and premature training termination. Furthermore, current evaluation methods rely heavily on downstream benchmarks (e.g., Needle-in-a-Haystack), which often fail to reflect the intrinsic convergence state and can lead to "deceptive saturation". In this paper, we present the first systematic investigation of LCCP learning dynamics using the industrial-grade Hunyuan-A13B (80B total parameters), tracking its evolution across a 200B-token training trajectory. Specifically, we propose a hierarchical framework to analyze LCCP dynamics across behavioral (supervised fine-tuning probing), probabilistic (perplexity), and mechanistic (attention patterns) levels. Our findings reveal: (1) Necessity of Massive Data Scaling: Training regimes of dozens of billions of tokens are insufficient for industrial-grade LLMs' LCCP (e.g., Hunyuan-A13B reaches saturation after training over 150B tokens). (2) Deceptive Saturation vs. Intrinsic Saturation: Traditional NIAH scores report "fake saturation" early, while our PPL-based analysis reveals continuous intrinsic improvements and correlates more strongly with downstream performance. (3) Mechanistic Monitoring for Training Stability: Retrieval heads act as efficient, low-resource training monitors, as their evolving attention scores reliably track LCCP progress and exhibit high correlation with SFT results. This work provides a comprehensive monitoring framework, evaluation system, and mechanistic interpretation for the LCCP of industrial-grade LLM.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。