arXiv:2607.07663cs.AI2026-07被引 10

梳理AI自我改进的四种路径,揭示自评估机制如何决定其安全边界。

Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops

论文配图:Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops
图 1 · 摘自论文原文
  • 按改进对象与闭环程度构建四维分类体系
  • 自评估信号强度决定改进效果,越强越接近人类判断
  • 指出研究方向设定是当前最需治理的关键瓶颈

AI系统正越来越多地参与自我改进:修正输出、部署中自适应调整、用自生成数据训练,甚至开展AI科研。现有文献使用‘自精炼’‘自奖励’‘自对弈’‘自演化’等术语,混同了不同目标。我们分析1250篇arXiv论文(2024-2026),从两个维度划分:系统改进的对象——部署行为、训练策略、评估器或研究过程本身——以及闭环程度(从人机协同到完全封闭)。该分类将有界自精炼(收敛、可评估、已工业应用)与开放递归自改进(RSI)区分开。后者受限于基础性要求、崩溃动态及算力约束。其核心特征是专门的自评估机制:每个改进循环本质上是声称某种信号可替代人类判断。我们梳理评估器设计空间——裁判、过程奖励模型、验证器、评分标准、元评估——并按验证强度排序,从形式化验证器(最强)到内在自评估(最弱)。实证表明,自改进能力与该层级一致,失败模式(自证循环、模型坍缩、多样性坍缩)源于对层级的违背。而‘研究方向设定’这一瓶颈使人类仍需介入,位于层级顶端。我们连接技术文献与递归自改进理论,回应前沿实验室闭环比对的安全治理问题,并指出‘治理级自改进度量’是当前最缺乏的研究领域。

原文摘要 · Abstract (English)

AI systems increasingly participate in their own improvement: revising their outputs, adapting their own harnesses during deployment, training on data they generate, and, increasingly, conducting AI research itself. This literature is described under a vocabulary ("self-refine," "self-reward," "self-play," "self-evolve") that conflates fundamentally different ambitions. We survey 1,250 arXiv papers (2024-2026) along two axes: what the system improves -- its behavior in deployment, its policy through training, its evaluator, or the research process itself -- and the degree of loop closure (human-in-the-loop to fully closed). The taxonomy separates bounded self-refinement -- convergent, evaluable, and already industrial practice -- from open-ended recursive self-improvement (RSI), which remains bounded by grounding requirements, collapse dynamics, and compute constraints on every measured axis. Its distinctive feature is a dedicated category for self-evaluation: every improvement loop is a claim that some signal can substitute for human judgment. We survey the evaluator design space -- judges, process reward models, verifiers, rubrics, meta-evaluation -- order the signals into a verification hierarchy from formal verifiers (strongest) to intrinsic self-assessment (weakest), and observe that demonstrated self-improvement strength tracks this hierarchy, that its failure modes (self-confirming loops, model collapse, diversity collapse) follow from its violations, and that the "research direction-setting" bottleneck keeping humans in the loop sits at the top of that hierarchy. We connect the technical literature to the theory of RSI limits and to the safety and governance questions raised by frontier-lab accounts of closing the loop, and identify governance-grade measurement of self-improvement as the field's most underpopulated niche.

AI自我改进自评估治理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。