arXiv:2603.29805cs.LGcs.AI2026-03被引 2

用量子化学方法预警深度学习训练中的相变,提前发现模型能力跃迁。

From Density Matrices to Phase Transitions in Deep Learning: Spectral Early Warnings and Interpretability

  • 基于双点约化密度矩阵,通过特征值统计捕捉训练过程中的相变信号。
  • 谱热容显示临界慢化,参与度比揭示系统重组维度,可提前预警二阶相变。
  • 主特征向量可直接解释,适合研究模型能力涌现的机制与适配场景。

现代人工智能研究的核心挑战之一是预测和理解模型在训练过程中涌现出的新能力。受量子化学反应分析方法的启发,我们提出「双点约化密度矩阵(2-datapoint reduced density matrix, 2RDM)」。该对象提供了一种计算高效的统一观测量,可用于追踪训练中的相变。通过滑动窗口分析2RDM的特征值统计,我们推导出两个互补信号:谱热容能通过临界慢化现象提前预警二阶相变;参与度比则揭示了底层重组的维度。令人惊讶的是,2RDM的前几个特征向量具有直接可解释性,便于深入研究相变的本质。我们在四种不同场景中验证了该方法的有效性:深度线性网络、归纳头形成、模型领悟(grokking)以及意外对齐(emergent misalignment)。最后,我们讨论了未来利用2RDM开展研究的方向。

原文摘要 · Abstract (English)

A key problem in the modern study of AI is predicting and understanding emergent capabilities in models during training. Inspired by methods for studying reactions in quantum chemistry, we present the ``2-datapoint reduced density matrix". We show that this object provides a computationally efficient, unified observable of phase transitions during training. By tracking the eigenvalue statistics of the 2RDM over a sliding window, we derive two complementary signals: the spectral heat capacity, which we prove provides early warning of second-order phase transitions via critical slowing down, and the participation ratio, which reveals the dimensionality of the underlying reorganization. Remarkably, the top eigenvectors of the 2RDM are directly interpretable making it straightforward to study the nature of the transitions. We validate across four distinct settings: deep linear networks, induction head formation, grokking, and emergent misalignment. We then discuss directions for future work using the 2RDM.

相变检测可解释性深度学习机制特征值分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。