无需环境模型,也能高效学习多样化技能。
Latent-Predictive Empowerment: Measuring Empowerment without a Simulator
- 用隐变量预测模型替代完整环境模型计算代理能力
- 在高维和随机环境中学到与主流方法相当的技能集
- 适合缺乏环境模型的真实场景,如机器人控制
代理能力有望帮助智能体学习大规模技能集,但目前尚难用于训练通用代理。现有方法通过最大化技能与状态间的互信息来学习多样技能,但需环境转移动态模型,这在高维、随机观测的真实场景中难以获得。本文提出潜变量预测能力(LPE),一种更实用的能力计算方法。LPE通过最大化一个可替代互信息的原理性目标,在仅需简单潜变量预测模型而非完整环境模拟器的前提下,学习大规模技能集。我们在多种场景中实证验证:(i)所提目标在高维观测和高度随机转移动态下,学到的技能集规模与依赖环境模型的领先方法相当;(ii)优于其他基于模型的能力方法。
原文摘要 · Abstract (English)
Empowerment has the potential to help agents learn large skillsets, but is not yet a scalable solution for training general-purpose agents. Recent empowerment methods learn diverse skillsets by maximizing the mutual information between skills and states; however, these approaches require a model of the transition dynamics, which can be challenging to learn in realistic settings with high-dimensional and stochastic observations. We present Latent-Predictive Empowerment (LPE), an algorithm that can compute empowerment in a more practical manner. LPE learns large skillsets by maximizing an objective that is a principled replacement for the mutual information between skills and states and that only requires a simpler latent-predictive model rather than a full simulator of the environment. We show empirically in a variety of settings--including ones with high-dimensional observations and highly stochastic transition dynamics--that our empowerment objective (i) learns similar-sized skillsets as the leading empowerment algorithm that assumes access to a model of the transition dynamics and (ii) outperforms other model-based approaches to empowerment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。