arXiv:2607.04277physics.soc-phcs.AI2026-07

提出大模型自我进化需突破内省能力阈值,否则会退化

Self-Reference in Large Language Models: The Introspection Threshold for Recursive Self-Improvement

论文配图:Self-Reference in Large Language Models: The Introspection Threshold for Recursive Self-Improvement
图 1 · 摘自论文原文
  • 用递归定理证明内省程序理论上存在
  • 实证发现现有模型仅具准内省,无法完全自访问
  • 指明架构改进路径及安全风险,适合关注AI自进化者

追求自演进人工智能的核心问题在于:何时自主自我改进是可持续的而非退化的?类比冯·诺伊曼对自复制自动机的复杂性阈值,我们认为大语言模型(LLMs)实现可持续递归自我改进需要一个功能等价物——内省,即系统模拟自身运行并针对修改目标的能力。基于克林第二递归定理,我们证明了此类内省程序在理论上存在。然而,实证回顾表明,尽管当前模型表现出准内省(如部分元认知),但因结构性瓶颈而未能达到真正内省:缺乏完整自访问、Transformer的前馈结构以及计算类约束导致无法进行不动点迭代。最后,我们概述了跨越这一复杂性阈值的架构路径,并讨论相关安全影响。

原文摘要 · Abstract (English)

The pursuit of self-evolving AI raises a critical question: when is autonomous self-improvement sustainable rather than degenerative? Drawing an analogy to von Neumann's complexity threshold for self-reproducing automata, we argue that sustainable recursive self-improvement in Large Language Models (LLMs) requires a functional analogue: introspection -- the system's capacity to simulate its own operations and target modifications. Grounded in Kleene's Second Recursion Theorem, we demonstrate the theoretical existence of such introspective programs. However, an empirical review reveals that while current LLMs exhibit quasi-introspection (e.g., partial metacognition), they fall short of true introspection due to structural bottlenecks: a lack of complete self-access, the feedforward nature of the Transformer, and computational class constraints that prevent fixed-point iteration. We conclude by outlining architectural paths to cross this complexity threshold and discussing the associated safety implications.

自进化内省机制大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。