用世界模型辅助蒸馏,让机器人从视觉中学会全身协同操作。
DreamMimic: Learning Visuomotor Whole-Body Loco-Manipulation via World Model

- 用世界模型学习预测性潜空间,提供多步监督信号
- 在OMOMO和BEHAVE上优于现有视觉基线方法
- 适合研究机器人视觉控制与接触丰富任务的学者
基于视觉的类人机器人全身运动操控面临部分可观测性、接触密集动力学及高维视觉输入下长时程行为学习困难等问题。我们提出DreamMimic框架,通过世界模型辅助蒸馏,将特权教师策略的知识迁移到视觉驱动的类人机器人控制器中。不同于传统Dreamer使用RSSM进行规划,我们将其重构为学习预测性潜动态,既作为表示空间,又作为动作条件的多步监督信号,并向学生策略暴露紧凑的预测特征以减少长期漂移。除标准的本体感知与视觉重建目标外,还引入额外的辅助预测头,用于预测特权状态、接触状态、物体状态和奖励估计,从而增强对人-物交互与任务进展的监督。进一步提出性能条件引导(PCG),一种基于奖励的自适应蒸馏调度机制,动态平衡教师指导与探索,防止在复杂视觉场景中过早降低教师影响或过度干扰。在OMOMO和BEHAVE数据集上的实验表明,该方法在无需部署时暴露在线特权交互状态的前提下,显著提升基于跟踪的全身操控性能。定性仿真验证了形态与模拟器变化的影响。结果表明,世界模型能有效稳定接触密集型类人行为的视觉策略蒸馏。
原文摘要 · Abstract (English)
Vision-based whole-body loco-manipulation on humanoid robots is challenging due to partial observability, contact-rich dynamics, and the difficulty of learning long-horizon behaviors from high-dimensional visual inputs. We present \href{https://github.com/DreamMimic/DreamMimic}{DreamMimic}, a framework that distills privileged teacher policies into vision-based humanoid controllers via world-model-assisted distillation. Instead of using a Dreamer-style RSSM for planning, we repurpose it to learn predictive latent dynamics that serve as both a representation space and an action-conditioned multi-step supervision signal, while exposing compact predictive features to the student policy to reduce long-term drift. Beyond standard reconstruction objectives for proprioceptive and visual observations, we add auxiliary prediction heads for privileged state, contact, object state, and reward estimation. These heads provide additional supervision related to agent--object interaction and task progress, encouraging the latent representation to retain signals that are useful for contact-rich loco-manipulation. We further introduce Performance-Conditioned Guidance (PCG), a reward-driven adaptive distillation schedule that computes performance scores for both teacher and student to dynamically balance guidance and exploration. PCG prevents both premature teacher annealing and excessive teacher interference in challenging visual settings. Experiments on OMOMO and BEHAVE show improved tracking-based loco-manipulation performance over strong vision-based baselines, without exposing online privileged interaction states to the student at deployment. Qualitative simulations further examine morphology and simulator changes. These results suggest that world models can provide a useful mechanism for stabilizing visual policy distillation in contact-rich humanoid behaviors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。