用可控信息生成理论统一智能与控制,让机器人自适应学习复杂行为。
Emergence of Physical Intelligence via Controllable Information Production
- 基于动力系统设计可控信息生成机制,避免人为设定变量偏差。
- 在标准机器人任务中超越已有方法,成功实现人形机器人翻正。
- 揭示价值函数结构与柯尔莫戈罗夫-西奈熵的关联,适合强化学习研究者。
内在动机(IM)旨在不依赖外部奖励的情况下训练智能体,使其仅通过与环境交互就涌现出有用行为。然而,主流的IM方法依赖于信息论量度及人为选定的变量,引入偏见且缺乏与动力学或最优控制(OC)的理论联系。本文提出可控信息生成(CIP),一种明确建立在动力系统和最优控制基础上的新式IM框架。CIP衡量智能体产生信息的速率,捕捉可控制的复杂性,无需外部知识或主观假设。该框架将IM与OC统一为单一理论体系,将物理智能定义为对信息生成的控制。同时,揭示了价值函数结构与柯尔莫戈罗夫-西奈熵之间的深层联系。CIP在标准机器人学习基准测试中持续优于先前方法,并解决了其无法完成的任务,包括人形机器人自我翻正。结果支持一个普遍原则:物理智能源于将系统驱动至可控混沌的边缘。
原文摘要 · Abstract (English)
Intrinsic Motivation (IM) aims to train agents without external rewards, enabling useful behavior to emerge from the agent's interaction with its environment alone. However, the dominant IM approaches rely on information-theoretic quantities with designer-chosen variables, introducing bias and lacking a principled connection to dynamics or optimal control (OC). We introduce Controllable Information Production (CIP), a new foundation for IM explicitly grounded in dynamical systems and OC. CIP measures the rate at which an agent produces information, capturing controllable complexity without external knowledge or bias. CIP unifies IM and OC into a single framework, formalizing physical intelligence as the control of information production. It further reveals connections between the structure of the value function and Kolmogorov-Sinai entropy. CIP consistently outperforms prior IM methods on standard benchmarks in robot learning and solves tasks they fail on, including humanoid self-righting. These results support a general organizing principle: physical intelligence emerges from driving systems toward the edge of controllable chaos.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。