大模型懂原理却不会算,根源在于执行路径与指令分离。
Comprehension Without Competence: Architectural Limits of LLMs in Symbolic Computation and Reasoning
- 发现模型存在指令与执行分离的架构缺陷
- 在数学和逻辑任务中表现不稳,即使提示完美也失败
- 适合研究模型本质局限与可解释性的人看
大型语言模型虽表面流畅,但在符号推理、算术准确性和逻辑一致性任务中系统性失败。本文通过受控实验与架构分析,揭示了‘理解’与‘能力’之间的持久鸿沟。模型常能正确表述原则,却无法可靠执行,问题不在知识获取,而在于计算执行。我们称此现象为‘计算分裂综合征’,即指令路径与执行路径在几何与功能上分离。这一核心缺陷贯穿数学运算与关系推理等多领域,解释了为何模型行为在理想提示下仍脆弱。我们认为,当前模型仅是强大的模式补全引擎,缺乏支撑原则性、组合式推理的架构支撑。研究明确了现有模型的能力边界,并呼吁未来模型引入元认知控制、原则提升与结构化执行机制。该诊断还表明,机械可解释性结果可能反映训练特定的模式协调,而非普遍计算原则;指令与执行路径的几何分离,暗示神经内省与机制分析的局限。
原文摘要 · Abstract (English)
Large Language Models (LLMs) display striking surface fluency yet systematically fail at tasks requiring symbolic reasoning, arithmetic accuracy, and logical consistency. This paper offers a structural diagnosis of such failures, revealing a persistent gap between \textit{comprehension} and \textit{competence}. Through controlled experiments and architectural analysis, we demonstrate that LLMs often articulate correct principles without reliably applying them--a failure rooted not in knowledge access, but in computational execution. We term this phenomenon the computational \textit{split-brain syndrome}, where instruction and action pathways are geometrically and functionally dissociated. This core limitation recurs across domains, from mathematical operations to relational inferences, and explains why model behavior remains brittle even under idealized prompting. We argue that LLMs function as powerful pattern completion engines, but lack the architectural scaffolding for principled, compositional reasoning. Our findings delineate the boundary of current LLM capabilities and motivate future models with metacognitive control, principle lifting, and structurally grounded execution. This diagnosis also clarifies why mechanistic interpretability findings may reflect training-specific pattern coordination rather than universal computational principles, and why the geometric separation between instruction and execution pathways suggests limitations in neural introspection and mechanistic analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。