提出新方法提升无监督强化学习技能的下游适应能力
Task Adaptation from Skills: Information Geometry, Disentanglement, and New Objectives for Unsupervised Reinforcement Learning

- 引入LSEPIN度量技能解耦性,优化技能多样性与可分性
- 用Wasserstein距离替代KL散度,使技能学习更利于任务迁移
- 理论证明新目标可发现更多最优初始策略,适合复杂任务预训练
无监督强化学习(URL)旨在为未见下游任务学习通用技能。互信息技能学习(MISL)通过最大化状态与技能间的互信息来实现,但缺乏充分的理论分析,例如其学习的技能能否有效初始化下游任务策略。本文的理论分析表明,技能的多样性和可分性对下游任务适配至关重要,而MISL并不保证这些性质。为此,我们提出新的解耦度量LSEPIN,并建立其与下游任务适应代价之间的信息几何联系。为进一步改善几何特性,我们研究用Wasserstein距离替代信息几何中的KL散度,推导出新目标WSEP,理论上更利于下游任务适配,且能发现比MISL更多的初始策略。最后,我们提出另一个基于Wasserstein距离的算法PWSEP,理论上可发现所有最优初始策略。
原文摘要 · Abstract (English)
Unsupervised reinforcement learning (URL) aims to learn general skills for unseen downstream tasks. Mutual Information Skill Learning (MISL) addresses URL by maximizing the mutual information between states and skills but lacks sufficient theoretical analysis, e.g., how well its learned skills can initialize a downstream task's policy. Our new theoretical analysis in this paper shows that the diversity and separability of learned skills are fundamentally critical to downstream task adaptation but MISL does not necessarily guarantee these properties. To complement MISL, we propose a novel disentanglement metric LSEPIN. Moreover, we build an information-geometric connection between LSEPIN and downstream task adaptation cost. For better geometric properties, we investigate a new strategy that replaces the KL divergence in information geometry with Wasserstein distance. We extend the geometric analysis to it, which leads to a novel skill-learning objective WSEP. It is theoretically justified to be helpful to downstream task adaptation and it is capable of discovering more initial policies for downstream tasks than MISL. We finally propose another Wasserstein distance-based algorithm PWSEP that can theoretically discover all optimal initial policies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。