通过技能多样性实现可识别表征,提升强化学习表征学习效果
Skill Learning via Policy Diversity Yields Identifiable Representations for Reinforcement Learning
- 利用技能多样性与对比成功特征方法,实现环境特征的可识别学习
- 在MuJoCo和DeepMind Control上验证了真实特征的线性恢复能力
- 为互信息目标设计提供理论依据,适合研究表征学习的学者
强化学习中的自监督特征学习与预训练方法常依赖信息论原理,即互信息技能学习(MISL)。这些方法旨在学习环境表征并激励探索。然而,表征与互信息参数化在MISL中的作用尚未有充分的理论理解。本文从可识别表征学习的角度,聚焦对比成功特征(CSF)方法,证明由于特征的内积参数化和技能多样性,CSF可在线性变换下严格恢复环境的真实特征。这是强化学习表征学习首个可识别性保证,同时解释了不同互信息目标的影响及熵正则化的副作用。我们在MuJoCo和DeepMind Control上实证验证了该结论,并展示了从状态和像素中均可严格恢复真实特征。
原文摘要 · Abstract (English)
Self-supervised feature learning and pretraining methods in reinforcement learning (RL) often rely on information-theoretic principles, termed mutual information skill learning (MISL). These methods aim to learn a representation of the environment while also incentivizing exploration thereof. However, the role of the representation and mutual information parametrization in MISL is not yet well understood theoretically. Our work investigates MISL through the lens of identifiable representation learning by focusing on the Contrastive Successor Features (CSF) method. We prove that CSF can provably recover the environment's ground-truth features up to a linear transformation due to the inner product parametrization of the features and skill diversity in a discriminative sense. This first identifiability guarantee for representation learning in RL also helps explain the implications of different mutual information objectives and the downsides of entropy regularizers. We empirically validate our claims in MuJoCo and DeepMind Control and show how CSF provably recovers the ground-truth features both from states and pixels.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。