arXiv:2511.11712cs.LGcs.AI2025-11

提出推理是状态空间中迭代算子收敛的机制,构建新模型实现76%准确率。

Reasoning: From Reflection to Solution

  • 将推理定义为状态空间中的迭代算子应用并收敛到固定点
  • 新模型OpenLM在OpenXOR任务上达76%准确率,远超现有LLM的0%
  • 适合研究可信赖推理机制与认知建模的学者

什么是推理?这一问题推动了从亚里士多德三段论到现代计算复杂性理论的千年哲学探索。在大语言模型于GSM8K(95%准确率)和HumanEval(90% pass@1)等基准上表现超越人类的当下,我们必须追问:这些系统是真正学会了推理,还是仅在模式匹配推理痕迹?本文主张一个具体答案:推理是状态空间中迭代算子应用并收敛到不动点的过程。这一定义不仅是哲学性的——它具有明确的架构意义,能解释当前系统的失败,并指明通往真实推理能力的道路。研究从一个谜题(OpenXOR)出发,经过理论分析(OpenOperator),最终提出可运行方案(OpenLM),在该任务上达到76%准确率,而当前顶尖大语言模型为0%。这并非批评现有系统,而是理解推理所需条件,并构建能够提供推理能力的架构。

原文摘要 · Abstract (English)

What is reasoning? This question has driven centuries of philosophical inquiry, from Aristotle's syllogisms to modern computational complexity theory. In the age of large language models achieving superhuman performance on benchmarks like GSM8K (95\% accuracy) and HumanEval (90\% pass@1), we must ask: have these systems learned to \emph{reason}, or have they learned to \emph{pattern-match over reasoning traces}? This paper argues for a specific answer: \textbf{reasoning is iterative operator application in state spaces, converging to fixed points}. This definition is not merely philosophical -- it has concrete architectural implications that explain both the failures of current systems and the path to genuine reasoning capabilities. Our investigation begins with a puzzle (OpenXOR), progresses through theory (OpenOperator), and culminates in a working solution (OpenLM) that achieves 76\% accuracy where state-of-the-art LLMs achieve 0\%. This is not about criticizing existing systems, but about \emph{understanding what reasoning requires} and \emph{building architectures that provide it}.

推理机制状态空间大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。