arXiv:2604.15350cs.LG2026-04被引 1

通过谱几何分析揭示Transformer推理的通用规律与可预测性。

The Spectral Geometry of Thought: Phase Transitions, Instruction Reversal, Token-Level Dynamics, and Perfect Correctness Prediction in How Transformers Reason

  • 用谱分析方法研究模型在推理时的隐藏状态动态变化。
  • 发现7种核心现象,包括推理时的谱压缩与正确性提前预测。
  • 适合对模型内部机制和推理过程感兴趣的开发者与研究者。

我们发现大型语言模型在推理与事实回忆时,其隐藏激活空间表现出显著的谱相变。通过对11个模型(涵盖5种架构:Qwen、Pythia、Phi、Llama、DeepSeek-R1)进行系统谱分析,识别出七种核心现象:(1) 推理谱压缩——11个模型中有9个在推理时α值显著更低(p < 0.05),且强模型效应更明显;(2) 指令微调谱反转——基础模型推理α < 事实α,而指令微调模型则相反;(3) 架构依赖生成分类——提示到回复的转变可分为扩张、压缩与平衡三类;(4) 谱缩放律——在4个Qwen基础模型中,α_推理 ∝ -0.074 ln N(R² = 0.46);(5) 逐标记谱级联——逐标记α追踪显示局部同步随层距指数衰减,推理任务中更弱;(6) 推理步骤谱标记——相变特征与推理步骤边界一致;(7) 谱正确性预测——仅靠谱α即可实现最高AUC=1.000(Qwen2.5-7B,后层),6模型平均AUC=0.893,可在最终答案生成前准确预测正确性。这些发现共同构建了变压器推理的完整谱理论,揭示思想的几何具有普遍方向、架构特异性动态,并可预测结果。

原文摘要 · Abstract (English)

We discover that large language models exhibit \emph{spectral phase transitions} in their hidden activation spaces when engaging in reasoning versus factual recall. Through systematic spectral analysis across \textbf{11 models} spanning \textbf{5 architecture families} (Qwen, Pythia, Phi, Llama, DeepSeek-R1), we identify \textbf{seven} core phenomena: (1)~\textbf{Reasoning Spectral Compression} -- 9/11 models show significantly lower $α$ for reasoning ($p < 0.05$), with larger effects in stronger models; (2)~\textbf{Instruction Tuning Spectral Reversal} -- base models show reasoning $α< $ factual $α$, while instruction-tuned models reverse this relationship; (3)~\textbf{Architecture-Dependent Generation Taxonomy} -- prompt-to-response shifts partition into expansion, compression, and equilibrium regimes; (4)~\textbf{Spectral Scaling Law} -- $α_\text{reasoning} \propto -0.074 \ln N$ across 4 Qwen base models ($R^2 = 0.46$); (5)~\textbf{Token-Level Spectral Cascade} -- per-token alpha tracking reveals local synchronization that decays exponentially with layer distance, and is weaker for reasoning than factual tasks; (6)~\textbf{Reasoning Step Spectral Punctuation} -- phase-transition signatures align with reasoning step boundaries; and (7)~\textbf{Spectral Correctness Prediction} -- spectral $α$ alone achieves AUC $= 1.000$ (Qwen2.5-7B, late layers) and mean AUC $= 0.893$ across 6 models in predicting correctness \emph{before} the final answer is generated. Together, these findings establish a comprehensive \emph{spectral theory of reasoning} in transformers, revealing that the geometry of thought is universal in direction, architecture-specific in dynamics, and predictive of outcome.

谱分析推理机制模型预测Transformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。