用语义掩码预测代替像素,让机器人学得更鲁棒
Mask World Model: Predicting What Matters for Robust Robot Policy Learning

- 用扩散模型预测物体掩码变化,过滤无关视觉噪声
- 在两个仿真基准上超越现有最先进方法,性能显著提升
- 适合追求高泛化能力的机器人控制研究者
基于大规模视频生成预训练的世界模型已成为通用机器人策略学习的有前景范式。然而,传统方法通常聚焦于高保真RGB视频预测,易过度拟合动态背景、光照变化等无关因素,降低模型泛化能力,导致控制策略不可靠且脆弱。为此,本文提出掩码世界模型(Mask World Model, MWM),采用视频扩散架构预测语义掩码演化而非像素。这一转变引入几何信息瓶颈,迫使模型捕捉关键物理动态与接触关系,同时过滤视觉噪声。我们无缝集成该掩码动态主干与基于扩散的策略头,实现端到端鲁棒控制。大量实验表明,MWM在LIBERO和RLBench仿真基准上显著优于当前最先进的基于RGB的世界模型。此外,真实世界实验与随机标记剪枝的鲁棒性评估显示,MWM展现出更强的泛化能力及对纹理信息丢失的鲁棒性。
原文摘要 · Abstract (English)
World models derived from large-scale video generative pre-training have emerged as a promising paradigm for generalist robot policy learning. However, standard approaches often focus on high-fidelity RGB video prediction, this can result in overfitting to irrelevant factors, such as dynamic backgrounds and illumination changes. These distractions reduce the model's ability to generalize, ultimately leading to unreliable and fragile control policies. To address this, we introduce the Mask World Model (MWM), which leverages video diffusion architectures to predict the evolution of semantic masks instead of pixels. This shift imposes a geometric information bottleneck, forcing the model to capture essential physical dynamics and contact relations while filtering out visual noise. We seamlessly integrate this mask dynamics backbone with a diffusion-based policy head to enable robust end-to-end control. Extensive evaluations demonstrate the superiority of MWM on the LIBERO and RLBench simulation benchmarks, significantly outperforming the state-of-the-art RGB-based world models. Furthermore, real-world experiments and robustness evaluation (via random token pruning) reveal that MWM exhibits superior generalization capabilities and robust resilience to texture information loss.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。