用干预复杂度定义智能的通用奖励,无需外部设定目标。
Intervention Complexity as a Canonical Reward and a Measure of Intelligence
- 以资源函数为依据,构建环境自生成的智能奖励
- 不同资源偏好对应可计算或不可计算的智能度量
- 适合研究通用智能、超级智能与预训练代理
Legg-Hutter通用智能度量将智能定义为在所有可计算环境中期望奖励的加权和,权重为简单性。但该框架依赖外部指定的奖励函数,引发奖励是否本质任意的问题。本文提出新度量——干预复杂度(intervention complexity),具备五个自然属性:环境衍生性、普遍性、最小性、敏感性和成就偏好。给定编码归纳偏置的资源函数ρ(如程序长度、执行时间或能量),ρ-干预复杂度即为通用奖励。由此形成一组由资源偏倚参数化的规范奖励,无需外部规范输入即可完成Legg-Hutter框架。进一步提出二维智能表征:代理能力(相对于最优神谕的表现)与学习效率(能力随经验提升的速度)。分离定理表明,资源偏倚的选择决定度量的可计算性:动作计数干预复杂度可在多项式时间内计算,而无神谕访问下的程序长度干预复杂度不可计算,二者差距精确刻画了学习的信息论内涵。对超智能与通用代理预训练具有启示意义。
原文摘要 · Abstract (English)
The Legg--Hutter universal intelligence measure provides a rigorous scalar assessment of general intelligence as expected reward across all computable environments, weighted by simplicity. However, the measure presupposes an externally specified reward function, raising the question of whether the reward primitive is inherently arbitrary or whether a canonical choice exists. We propose a new measure, called intervention complexity, that has five natural properties: environment-derivedness, universality, minimality, sensitivity, and achievement preference. Given a resource function rho encoding an inductive bias (such as program length, execution time, or energy), rho-intervention complexity is a universal reward. The result yields a family of canonical rewards indexed by resource bias, providing a principled completion of the Legg--Hutter framework that does not require external normative input. We further propose a two-dimensional characterisation of intelligence: agent competence (how well the agent performs relative to the oracle optimum) and learning efficiency (how quickly this competence improves with experience). A separation theorem establishes that the choice of resource bias determines the computability of the resulting measure: action-count IC is computable in polynomial time, while program-length IC without oracle access is uncomputable, with the gap between oracle and bare IC precisely quantifying the information-theoretic content of learning. We discuss implications for superintelligence and for pre-training universal agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。