用统计物理方法解析语言模型的推理动态,揭示其内在思维模式。
A Statistical Physics of Language Model Reasoning
- 将推理过程建模为低维流形上的随机动力系统,捕捉思维演变轨迹。
- 40维投影解释约50%的推理状态方差,发现四种隐含推理阶段。
- 可低成本模拟推理失败,适合研究模型错误与关键转折点。
Transformer语言模型展现出难以解析的涌现式推理能力。本文提出一种统计物理框架,用于描述连续时间链式思考推理动态。将句子级隐藏状态轨迹建模为低维流形上的随机动力系统,通过潜在状态切换捕捉多样化的推理阶段,包括错位状态或失败情况。对8个模型、7个基准的数据进行实证分析显示,40维投影(在方差保留与可行性间平衡)可解释约50%的方差。我们识别出四种潜在推理模式,并构建了可验证的隐马尔可夫状态动态系统(SLDS)来捕捉这些特征。该框架支持低成本推理模拟,为研究和预测关键转变(如错位状态或其他模型失败)提供了工具。
原文摘要 · Abstract (English)
Transformer LMs show emergent reasoning that resists mechanistic understanding. We offer a statistical physics framework for continuous-time chain-of-thought reasoning dynamics. We model sentence-level hidden state trajectories as a stochastic dynamical system on a lower-dimensional manifold. This drift-diffusion system uses latent regime switching to capture diverse reasoning phases, including misaligned states or failures. Empirical trajectories (8 models, 7 benchmarks) show a rank-40 projection (balancing variance capture and feasibility) explains ~50% variance. We find four latent reasoning regimes. An SLDS model is formulated and validated to capture these features. The framework enables low-cost reasoning simulation, offering tools to study and predict critical transitions like misaligned states or other LM failures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。