提出可量化大模型训练中分布退化的谱分析方法
SIGMA: Scalable Spectral Insights for LLM Model Collapse
- 基于嵌入矩阵谱特性构建统一评估框架
- 能捕捉训练过程中表示空间收缩的退化过程
- 适用于大规模模型,可实时监控递归训练健康度
大规模语言模型(LLM)广泛采用合成数据训练,但由此引发的“模型坍塌”问题日益突出——在模型生成内容上反复训练会导致分布方差缩小和表征质量下降。尽管坍塌现象已被广泛观察,但在高维空间中量化与预测其发生仍缺乏严谨方法。本文提出SIGMA(Spectral Inequalities for Gram Matrix Analysis),一种基于嵌入Gram矩阵谱特性的统一评估框架。通过推导并利用矩阵谱的确定性与随机界,SIGMA提供了数学严谨的指标来追踪表示空间的收缩。其随机形式支持对大规模基础模型进行可扩展估计,避免全特征分解的计算瓶颈。实验表明,SIGMA能有效捕捉向退化状态演进的过程,兼具理论深度与实际应用价值,为递归训练管道的健康监测提供可靠工具。
原文摘要 · Abstract (English)
The rapid adoption of synthetic data for training Large Language Models (LLMs) has introduced the technical challenge of "model collapse"-a degenerative process where recursive training on model-generated content leads to a contraction of distributional variance and representational quality. While the phenomenology of collapse is increasingly evident, rigorous methods to quantify and predict its onset in high-dimensional spaces remain elusive. In this paper, we introduce SIGMA (Spectral Inequalities for Gram Matrix Analysis), a unified framework that benchmarks model collapse through the spectral lens of the embedding Gram matrix. By deriving and utilizing deterministic and stochastic bounds on the matrix's spectrum, SIGMA provides a mathematically grounded metric to track the contraction of the representation space. Crucially, our stochastic formulation enables scalable estimation of these bounds, making the framework applicable to large-scale foundation models where full eigendecomposition is intractable. We demonstrate that SIGMA effectively captures the transition towards degenerate states, offering both theoretical insights into the mechanics of collapse and a practical, scalable tool for monitoring the health of recursive training pipelines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。