arXiv:2602.06413cs.AI2026-02

发现自回归推理有内在稳定性极限,长链条推理必须分段才能稳定。

Intrinsic Stability Limits of Autoregressive Reasoning: Structural Consequences for Long-Horizon Execution

  • 自回归生成在长链条中会指数级衰减决策优势,导致推理崩溃。
  • 实验证明合成任务和TextWorld中存在符合理论预测的性能突降。
  • 提示未来推理系统需转向分段结构而非单纯扩大模型规模。

大型语言模型虽具强大推理能力,但在长链条任务中表现常急剧下降,存在系统性失效。传统解释多归因于任务复杂度,如组合爆炸或长期信用分配难题。本文提出,即使在无语义歧义、单路径线性任务中,自回归执行也存在内在稳定性极限。我们指出,长链条推理的根本约束来自生成过程本身的不稳定性,而非仅由搜索或任务复杂性导致,将长链条推理重新定义为结构治理问题。推导出定理A:单路径自回归推理中的决策优势随执行长度呈指数衰减,确立了可维持推理链的理论上限。这暗示稳定长链条推理需离散分段,自然催生类似有向无环图(DAG)的执行结构。在合成环境与真实TextWorld任务中的实证研究显示,性能突降现象与理论预测一致。研究揭示了长链条推理失败的动力学机制,表明纯自回归架构难以维持长期一致性。此外,短链条评估可能掩盖结构性不稳定性,提示未来推理系统应从规模扩展转向结构治理。

原文摘要 · Abstract (English)

Large language models (LLMs) demonstrate remarkable reasoning capabilities, yet their performance often deteriorates sharply in long-horizon tasks, exhibiting systematic breakdown beyond certain scales. Conventional explanations primarily attribute this phenomenon to task complexity, such as combinatorial search explosion or long-term credit assignment challenges. In this work, we argue that these explanations are incomplete: even in linear, unbranched tasks without semantic ambiguity, autoregressive execution is subject to an intrinsic stability limit. We propose that the fundamental constraint on long-horizon reasoning arises from process-level instability in autoregressive generation rather than solely from search or task complexity, reframing long-horizon reasoning as a problem of structural governance. We derive Theorem~A, showing that decision advantage in single-path autoregressive reasoning decays exponentially with execution length, imposing a fundamental bound on maintainable reasoning chains. This result implies a structural consequence: stable long-horizon reasoning requires discrete segmentation, naturally inducing graph-like execution structures such as directed acyclic graphs (DAGs). Empirical studies in both synthetic environments and real TextWorld tasks reveal observable performance cliffs consistent with theoretical predictions. Our findings provide a dynamical perspective on long-horizon reasoning failure and suggest new limitations on maintaining long-term coherence under purely autoregressive architectures. Furthermore, we highlight that short-horizon evaluation protocols may obscure structural instability, indicating a potential shift from scaling toward structured governance in future reasoning systems.

自回归推理稳定性长链条结构治理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。