arXiv:2603.10384cs.AI2026-03中稿 · ICML被引 2

用几何轨迹分析大模型推理,区分真思考与胡扯。

Beyond Scalars: Evaluating and Understanding LLM Reasoning via Geometric Progress and Stability

  • 将推理过程拆解为进展和稳定性两个几何维度
  • 正确推理呈现高进展、稳定轨迹,幻觉则低进展且波动大
  • 可解释性强,适合研究模型内在思维机制的学者

通过标量概率评估大模型可靠性往往无法捕捉推理的结构动态。我们提出TRACED框架,基于理论支撑的几何运动学评估推理质量。将推理轨迹分解为进展(位移)与稳定性(曲率),揭示显著拓扑差异:正确推理表现为高进展、稳定轨迹,而幻觉则呈现低进展、不稳定的模式(位移停滞且曲率波动剧烈)。利用这些特征,我们的概率框架在多个基准上实现竞争力表现和更强鲁棒性。关键的是,TRACED通过将高曲率映射到“犹豫环”,位移映射到“确定性累积”,实现了几何与认知的桥梁,为解码机器思维内部动态提供了物理视角。

原文摘要 · Abstract (English)

Evaluating LLM reliability via scalar probabilities often fails to capture the structural dynamics of reasoning. We introduce TRACED, a framework that assesses reasoning quality through theoretically grounded geometric kinematics. By decomposing reasoning traces into Progress (displacement) and Stability (curvature), we reveal a distinct topological divergence: correct reasoning manifests as high-progress, stable trajectories, whereas hallucinations are characterized by low-progress, unstable patterns (stalled displacement with high curvature fluctuations). Leveraging these signatures, our probabilistic framework achieves competitive performance and superior robustness across diverse benchmarks. Crucially, TRACED bridges geometry and cognition by mapping high curvature to ''Hesitation Loops'' and displacement to ''Certainty Accumulation'', offering a physical lens to decode the internal dynamics of machine thought.

大模型推理几何分析可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。