提出GenDa框架,提升无监督强化学习的泛化与数据效率。
Learning Generalizable Skill Policy with Data-Efficient Unsupervised RL

- 通过技能重标注缓解语义不稳定性,提升预训练数据效率。
- 设计互补信息瓶颈,使策略聚焦自身特征,适应分布偏移。
- 适用于需要强泛化能力的下游控制任务,尤其适合数据稀缺场景。
无监督强化学习(URL)旨在无需外部奖励的情况下预训练可扩展的技能条件策略,为下游控制任务奠定基础。尽管近期取得进展,我们指出当前离策略URL方法存在两个被忽视的关键瓶颈:(1) 技能语义非平稳性,(2) 泛化能力脆弱。为此,我们提出GenDa(通用高效智能体)统一框架,以实现鲁棒的无监督强化学习。首先,引入技能重标注机制,缓解非平稳性,显著提升预训练数据效率。其次,提出互补信息瓶颈(CIB),促使学习到的技能策略聚焦于自身中心特征,增强对下游任务分布偏移的鲁棒性。通过多项实验验证,GenDa显著提升了URL的可扩展性,具备更优的泛化能力和数据效率。代码与视频见 https://ihatebroccoli.github.io/official-GenDa。
原文摘要 · Abstract (English)
Unsupervised Reinforcement Learning (URL) aims to pre-train scalable, skill-conditioned policies without extrinsic rewards, serving as a foundation for downstream control tasks. Despite recent progress, we argue that current off-policy URL methods are limited by two critical, overlooked bottlenecks: (1) non-stationary skill semantics and (2) brittle generalization. To address these challenges, we propose GenDa (Generalizable Data-efficient Agent), a unified framework for robust unsupervised reinforcement learning. First, we introduce a skill relabeling mechanism to mitigate non-stationarity and significantly improve data efficiency for pre-training. Second, we propose a Complementary Information Bottleneck (CIB), encouraging the learned skill policy to focus on ego-centric features and become robust to distribution shifts for downstream tasks. Through various experiments, we demonstrate that GenDa significantly enhances the scalability of URL with superior generalizability and data efficiency. Our code and videos are available at https://ihatebroccoli.github.io/official-GenDa.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。