提出可量化智能体的理论框架,揭示其自发出现的可能性。
Active Inference Agency Formalization, Metrics, and Convergence Assessments
- 用主动推断构建智能体:平衡好奇与控制力实现自指
- 智能体函数光滑凸,稀疏环境中对数收敛,易自发产生
- 定义新指标测量智能体程度,可用于检测危险内生优化
本文针对人工智能安全中的嵌套优化问题,提出智能体的形式化定义与分析框架。智能体被建模为持续累积经验的连续表示,通过动态平衡好奇心(最小化预测误差以维持不可计算性与新颖性)和赋能(最大化控制通道的信息容量以保证主体性与目标导向性)实现自指。实证表明该主动推断模型能有效解释经典工具性目标,如自我保存与资源获取。分析显示所提智能体函数光滑且凸,具备良好优化性质。尽管智能体函数在抽象函数空间中占比极小,但在稀疏环境中表现出对数收敛,暗示现代大规模模型训练中智能体自发出现概率很高。为此,论文提出基于标准化奖励空间(STARC)中行为等价距离的智能体度量方法,可量化系统与理想智能体目标的接近程度,为分类和检测嵌套优化器提供可靠工具。
原文摘要 · Abstract (English)
This paper addresses the critical challenge of mesa-optimization in AI safety by providing a formal definition of agency and a framework for its analysis. Agency is conceptualized as a Continuous Representation of accumulated experience that achieves autopoiesis through a dynamic balance between curiosity (minimizing prediction error to ensure non-computability and novelty) and empowerment (maximizing the control channel's information capacity to ensure subjectivity and goal-directedness). Empirical evidence suggests that this active inference-based model successfully accounts for classical instrumental goals, such as self-preservation and resource acquisition. The analysis demonstrates that the proposed agency function is smooth and convex, possessing favorable properties for optimization. While agentic functions occupy a vanishingly small fraction of the total abstract function space, they exhibit logarithmic convergence in sparse environments. This suggests a high probability for the spontaneous emergence of agency during the training of modern, large-scale models. To quantify the degree of agency, the paper introduces a metric based on the distance between the behavioral equivalents of a given system and an "ideal" agentic function within the space of canonicalized rewards (STARC). This formalization provides a concrete apparatus for classifying and detecting mesa-optimizers by measuring their proximity to an ideal agentic objective, offering a robust tool for analyzing and identifying undesirable inner optimization in complex AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。