发现大模型推理时隐藏的几何脉冲,让概念分离更易计算。
Emergent Manifold Separability during Reasoning in Large Language Models
- 用流形容量理论分析推理过程中的表示几何变化。
- 推理前概念流形短暂分离成线性可分子空间,后迅速压缩。
- 揭示动态流形管理机制,适合研究模型推理机理者阅读。
链式思维提示显著提升大语言模型的推理能力,但其潜在表征几何的时间动态仍不清晰。本文通过流形容量理论(MCT)研究两个组合推理任务:一个可控的布尔逻辑树,支持深度机制分析;一个自然语言资格判断任务,要求从文本中提取属性、与阈值比较,并通过固定评估树组合局部决策。MCT 可在无需探针训练的情况下量化潜在表示的线性可分性。在两个任务及多个开源模型上,推理表现为瞬态几何脉冲:概念流形在计算前瞬间解缠为线性可分子空间,随后快速压缩。该行为与标准线性探针准确率长期保持高位形成对比,表明信息可检索与信息几何就绪存在根本差异。我们将其解释为动态流形管理,即模型动态调节表征容量,以优化推理链中残差流的带宽。
原文摘要 · Abstract (English)
Chain-of-Thought (CoT) prompting significantly improves reasoning in Large Language Models, yet the temporal dynamics of the underlying representation geometry remain poorly understood. We investigate these dynamics by applying Manifold Capacity Theory (MCT) to two compositional reasoning tasks: a controlled Boolean logic tree that supports deep mechanistic analysis, and a natural-language eligibility task in which the model has to extract attributes from prose, compare them to thresholds, and compose the local decisions through a fixed evaluation tree. MCT lets us quantify the linear separability of latent representations without the confounding factors of probe training. On both tasks, and across several open-weight models, reasoning manifests as a transient geometric pulse: concept manifolds are untangled into linearly separable subspaces immediately prior to computation and rapidly compressed thereafter. This behavior diverges from standard linear probe accuracy, which remains high long after computation, suggesting a fundamental distinction between information that is merely retrievable and information that is geometrically prepared for processing. We interpret this phenomenon as Dynamic Manifold Management, a mechanism where the model dynamically modulates representational capacity to optimize the bandwidth of the residual stream throughout the reasoning chain.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。