让AI自己改进学习方法,还能控制风险。
Emotion-Gradient Metacognitive RSI (Part I): Theoretical Foundations and Single-Agent Architecture
- 用情绪和信心驱动奖励,让智能体自我优化学习算法。
- 提出意义密度与转换效率,量化语义学习效果。
- 适合研究安全可控的通用人工智能架构的人看。
我们提出情感梯度元认知递归自改进(EG-MRSI)框架,将内省式元认知、基于情绪的内在动机与递归自修改整合为统一理论体系。该框架在形式化风险约束下可自主重写自身学习算法。基于噪声到意义递归自改进(N2M-RSI)基础,EG-MRSI引入由置信度、误差、新颖性与累积成功驱动的可微分内在奖励函数,调控元认知映射与受可证明安全机制约束的自修改算子。本文正式定义初始代理配置、情感梯度动态及自改进触发条件,并推导出与强化学习兼容的优化目标,引导智能体发展轨迹。引入意义密度与意义转换效率作为语义学习的可量化指标,弥合内部结构与预测信息量之间的鸿沟。本部分建立EG-MRSI单智能体理论基础。后续部分将扩展至安全证书与回滚协议(第二部分)、集体智能机制(第三部分)以及热力学与计算极限等可行性约束(第四部分)。整个系列为开放且安全的通用人工智能提供严谨、可扩展的基础。
原文摘要 · Abstract (English)
We present the Emotion-Gradient Metacognitive Recursive Self-Improvement (EG-MRSI) framework, a novel architecture that integrates introspective metacognition, emotion-based intrinsic motivation, and recursive self-modification into a unified theoretical system. The framework is explicitly capable of overwriting its own learning algorithm under formally bounded risk. Building upon the Noise-to-Meaning RSI (N2M-RSI) foundation, EG-MRSI introduces a differentiable intrinsic reward function driven by confidence, error, novelty, and cumulative success. This signal regulates both a metacognitive mapping and a self-modification operator constrained by provable safety mechanisms. We formally define the initial agent configuration, emotion-gradient dynamics, and RSI trigger conditions, and derive a reinforcement-compatible optimization objective that guides the agent's development trajectory. Meaning Density and Meaning Conversion Efficiency are introduced as quantifiable metrics of semantic learning, closing the gap between internal structure and predictive informativeness. This Part I paper establishes the single-agent theoretical foundations of EG-MRSI. Future parts will extend this framework to include safety certificates and rollback protocols (Part II), collective intelligence mechanisms (Part III), and feasibility constraints including thermodynamic and computational limits (Part IV). Together, the EG-MRSI series provides a rigorous, extensible foundation for open-ended and safe AGI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。