用随机网络蒸馏构建联合分布奖励,提升世界模型在线模仿学习的稳定性。
Coupled Distributional Random Expert Distillation for World Model Online Imitation Learning
- 基于潜空间中专家与行为分布的联合估计设计奖励模型。
- 在DMControl、Meta-World等基准上实现专家级表现且训练更稳定。
- 适合需要稳定模仿学习的机器人控制与操作任务场景。
模仿学习(IL)已在机器人、自动驾驶和医疗等领域取得显著进展,使智能体能从专家示范中学习复杂行为。然而,现有方法在世界模型框架中依赖对抗性奖励或价值函数时,常面临不稳定性问题。本文提出一种新型在线模仿学习方法,通过基于随机网络蒸馏(RND)的密度估计构建奖励模型,联合估计世界模型潜空间中的专家分布与行为分布。我们在多种基准上进行评估,包括DMControl、Meta-World和ManiSkill2,结果表明该方法在行走与操作任务中均实现专家级性能,同时相比对抗方法显著提升了训练稳定性。
原文摘要 · Abstract (English)
Imitation Learning (IL) has achieved remarkable success across various domains, including robotics, autonomous driving, and healthcare, by enabling agents to learn complex behaviors from expert demonstrations. However, existing IL methods often face instability challenges, particularly when relying on adversarial reward or value formulations in world model frameworks. In this work, we propose a novel approach to online imitation learning that addresses these limitations through a reward model based on random network distillation (RND) for density estimation. Our reward model is built on the joint estimation of expert and behavioral distributions within the latent space of the world model. We evaluate our method across diverse benchmarks, including DMControl, Meta-World, and ManiSkill2, showcasing its ability to deliver stable performance and achieve expert-level results in both locomotion and manipulation tasks. Our approach demonstrates improved stability over adversarial methods while maintaining expert-level performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。