揭示自改进智能体的统计极限,证明能力无界会导致可学习任务变不可学。
On The Statistical Limits of Self-Improving Agents
- 将自修改分解为五个维度,建立分析框架
- 能力无界时,理性自我优化会使可学习任务失效
- 提出双闸门机制,保障学习能力与泛化率
我们构建了一个学习理论框架,用于分析自改进智能体,将其自修改行为分解为五个维度。在此框架下,我们证明了一个严格边界:在标准独立同分布假设下,分布无关的PAC可学习性得以保持,当且仅当策略可达家族保持统一容量有界。若可达容量可无限增长,效用理性的自我改变会使原本可学习的任务变得不可学习。我们进一步提出一种简单的双闸门防护机制——验证-改进要求结合容量上限——能维持该边界,并获得标准VC率保证。更广泛的启示是,自修改不仅需受目标约束,还需满足保障学习统计前提的结构条件。随着人工智能系统日益智能和自主,本框架为自改进的统计理论奠定了基础。
原文摘要 · Abstract (English)
We develop a learning-theoretic framework for analyzing self-improving agents by decomposing self-modification into five axes. Within this framework, we prove a sharp boundary: under standard i.i.d. assumptions, distribution-free PAC learnability is preserved if and only if the policy-reachable family remains uniformly capacity-bounded. If reachable capacity can grow without bound, utility-rational self-changes can make learnable tasks unlearnable. We further introduce a simple Two-Gate guardrail -- a validation-improvement requirement plus a capacity cap -- that preserves this boundary and yields standard VC-rate guarantees. The broader implication is that self-modification must be constrained not only by objectives, but also by structural conditions that preserve the statistical prerequisites for learning. As AI systems become increasingly intelligent and autonomous, this framework provides a foundation for the statistical theory of self-improvement.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。