arXiv:2605.04330cs.AIcs.CC2026-05

深度Transformer通过隐式推理逼近显式思维链表现。

The Scaling Properties of Implicit Deductive Reasoning in Transformers

论文配图:The Scaling Properties of Implicit Deductive Reasoning in Transformers
图 1 · 摘自论文原文
  • 用双向前缀掩码构建深度模型,实现隐式推理
  • 在不同图结构和问题宽度下接近显式思维链效果
  • 适合研究大模型推理机制的学者参考

我们系统研究了深度受限的Transformer在命题逻辑(Horn clauses)上的隐式演绎推理缩放特性。通过解耦可证明性与虚假特征,并强制算法对齐,发现足够深的模型配合双向前缀掩码时,其隐式推理性能可覆盖多种图拓扑和问题宽度,接近显式思维链(CoT)水平,但深度外推仍需显式CoT支持。

原文摘要 · Abstract (English)

We investigate the scaling properties of implicit deductive reasoning over Horn clauses in depth-bounded Transformers. By systematically decorrelating provability from spurious features and enforcing algorithmic alignment, we find that in sufficiently deep models with a bidirectional prefix mask, implicit reasoning approaches explicit CoT performance across graph topologies and problem widths, though CoT remains necessary for depth extrapolation.

Transformer推理缩放定律

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。