arXiv:2609.09030cs.AIcs.CL2026-09

用动态轨迹分析大模型推理过程,揭示答案演变背后的机制。

Answer-Distribution Trajectories: A Stochastic-Dynamics View of LLM Reasoning

  • 构建答案分布轨迹,追踪推理中所有可能答案的变化
  • 相同终点和熵值下,不同模型有截然不同的推理路径
  • 适合研究推理失败原因或优化训练策略的从业者

链式思维推理为模型从输入到最终答案提供了结构化计算过程,但通常仅通过终点准确率评估,忽略了推理路径。现有工作使用熵谱追踪不确定性演变,但无法揭示导致不确定性的竞争假设。本文提出答案分布轨迹,一种受随机动力学启发的表示方法,追踪推理过程中模型对答案的完整预测分布。相比终点结果和熵摘要,该表示更精细,可刻画探索、修正、演进与确定等推理动态阶段,并区分推理成功与失败的不同机制。在16个开源语言模型和4个推理基准上,我们发现具有相同终点和相似熵谱的轨迹,其推理动态可能差异显著。同时观察到模型内与跨任务间存在明显动态差异,不同目标偏好不同动态特征。此外,训练和推理选择会系统性重塑这些动态。结果表明,答案分布轨迹为分析与评估大模型推理动态提供了一个丰富框架。

原文摘要 · Abstract (English)

Chain-of-thought reasoning provides a structured computation between a model's input and final answer. Yet it is often evaluated through endpoint accuracy, which ignores the path taken to reach that answer. An emerging line of work addresses this limitation using entropy profiles, which track how uncertainty evolves over the reasoning process but do not reveal which competing hypotheses account for that uncertainty. We introduce answer-distribution trajectories, a stochastic-dynamics-inspired representation that tracks the model's full predictive distribution over answers as reasoning unfolds. As a strictly finer representation than endpoint and entropy summaries, answer-distribution trajectories enable us to characterize a trace through a dynamical reasoning profile spanning exploration, revision, motion, and commitment, and to distinguish different dynamical mechanisms of reasoning success and failure. Across sixteen open-weight language models and four reasoning benchmarks, we show that traces with the same endpoint and similar entropy profiles can exhibit substantially different reasoning dynamics. We further find substantial variation in these dynamics both within and across models and tasks, with different objectives favoring different dynamical profiles. Additionally, we show that training and inference choices systematically reshape these profiles. Our results suggest that answer-distribution trajectories provide a rich framework for analysing and evaluating the dynamics of LLM reasoning.

大模型推理动态分析链式思维

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。