arXiv:2602.00969cs.LG2026-02

低精度训练时,量化会破坏语义嵌入的谱结构,导致模型崩溃。

On the Spectral Flattening of Quantized Embeddings

  • 从齐普夫定律与随机矩阵理论出发,证明语义嵌入的幂律谱是必需的。
  • 均匀量化引入噪声底限,使谱尾被严重截断,稳定秩显著上升。
  • 揭示了低比特优化中保持谱保真度的重要性,适合关注模型稳定性研究者。

在超低精度下训练大语言模型(LLMs)面临严重不稳定性,其根源在于离散量化约束与语言数据固有的重尾谱特性之间的冲突。通过将齐普夫统计与随机矩阵理论相联系,我们证明嵌入奇异值谱的幂律衰减是语义编码的基本要求。理论上推导出,均匀量化引入一个噪声底限,导致谱尾被不成比例地截断,从而引发谱平坦化,并严格证明表示的稳定秩显著增加。在GPT-2和TinyLlama等多种架构上的实证验证表明,这种几何退化会引发表征崩溃。本工作不仅量化了LLMs对谱敏感性的程度,还确立谱保真度为稳定低比特优化的必要条件。

原文摘要 · Abstract (English)

Training Large Language Models (LLMs) at ultra-low precision is critically impeded by instability rooted in the conflict between discrete quantization constraints and the intrinsic heavy-tailed spectral nature of linguistic data. By formalizing the connection between Zipfian statistics and random matrix theory, we prove that the power-law decay in the singular value spectra of embeddings is a fundamental requisite for semantic encoding. We derive theoretical bounds showing that uniform quantization introduces a noise floor that disproportionately truncates this spectral tail, which induces spectral flattening and a strictly provable increase in the stable rank of representations. Empirical validation across diverse architectures including GPT-2 and TinyLlama corroborates that this geometric degradation precipitates representational collapse. This work not only quantifies the spectral sensitivity of LLMs but also establishes spectral fidelity as a necessary condition for stable low-bit optimization.

低精度训练谱分析嵌入质量模型稳定性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。