arXiv:2606.09663cs.AI2026-06

提出可复现的元AI自设计框架,验证了系统自我迭代提升能力。

From 0-to-1 to 1-to-N: Reproducible Engineering Evidence for MetaAI Recursive Self-Design

  • 构建四维度评估框架,定义元级修改与递归改进机制
  • DGM系统经80轮迭代,代码任务准确率从20%提至50%
  • 开源可复现协议MetaAI-Mini,适合研究自进化AI的开发者

递归自设计指人工智能辅助修改自身构建、评估与优化机制。本文将MetaAI视为一种由人类启动、由AI扩展的开发模式,其设计空间本身成为可修改目标。提出包含四项标准的操作性证据框架:可检视的目标系统、元层级修改器、反馈驱动的选择机制以及递归延续性。将公开系统如达尔文哥德尔机(DGM)、STOP、哥德尔代理和ShinkaEvolve映射至该框架。DGM提供当前最直接证据:在SWE-bench Verified上准确率从20%提升至50%,在完整Polyglot上从14.2%升至30.7%,经过80轮迭代;消融实验表明开放探索与自我改进均起作用。最后,提出MetaAI-Mini,一个基于HumanEval的可复现协议与代码库。因未包含完整模型运行,该版本仅作为协议发布。

原文摘要 · Abstract (English)

Recursive self-design refers to AI-assisted modification of the mechanisms by which an AI system is built, evaluated, and improved. This paper treats MetaAI not as a mature paradigm, but as a working term for a human-seeded, AI-expanded development pattern in which the design space itself becomes a target of modification. We propose an operational evidence framework with four criteria: inspectable target system, meta-level modifier, feedback-directed selection, and recursive continuation. We then map public systems, including Darwin Goedel Machine (DGM), STOP, Goedel Agent, and ShinkaEvolve, against these criteria. DGM provides the most direct currently reported evidence: its published results show improvement from 20% to 50% on SWE-bench Verified and from 14.2% to 30.7% on full Polyglot after 80 iterations, with ablations suggesting that both open-ended exploration and self-improvement contribute. Finally, we provide MetaAI-Mini, a reproducible HumanEval-based protocol and codebase. Because no completed model run is included in this build, MetaAI-Mini is reported as a protocol rather than as an experimental result.

元AI自进化可复现代码生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。