用信息论定义持续学习,提出可衡量的开放性标准。
An Information-Theoretic Definition for Open-Ended Learning
- 以比特等价量度奖励增长所需信息,定义开放性
- 证明经典多臂赌博机非开放性,构造出开放性环境
- 设计算法实现该环境下持续学习,适合强化学习研究者
越来越多的研究表明,能在开放环境中持续拓展能力的智能体具有巨大潜力。然而,目前尚无对开放性的清晰定义或探索机制的理论指导。本文基于新概念‘比特等价’(bit-equivalent),提出一种信息论定义:若智能体能实现比特等价的线性增长,则环境为开放性。我们证明经典多臂赌博机环境不具开放性,并构造出一个具备开放性的赌博机环境。此外,还提出一种可在该环境中实现开放性学习的算法。
原文摘要 · Abstract (English)
A growing body of work points to the great promise of AI systems that can continually expand their capabilities as they operate in an open-ended environment. But yet there is no coherent definition of open-endedness or theory about how an agent ought to explore an open-ended environment. We introduce an information-theoretic definition based on a new concept -- the ${\textit bit-equivalent}$ -- which quantifies the information required to attain each level of expected reward. We consider an environment to be open-ended if an agent can attain linear growth in the bit-equivalent. We establish that classical bandit environments are not open-ended and formulate a bandit environment that is. We also introduce an algorithm that achieves open-ended learning in this environment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。