arXiv:2509.21797cs.CV2025-09被引 4

融合隐空间与像素特征,提升机器人操作规划精度与泛化能力

MoWM: Mixture-of-World-Models for Embodied Planning via Latent-to-Pixel Feature Modulation

  • 用混合世界模型融合运动感知隐空间与像素级特征
  • 在CALVIN和真实场景任务中达成最优成功率
  • 适合关注机器人视觉动作规划的研究者

具身行动规划是机器人领域的核心挑战,需从视觉观测和语言指令生成精确动作。尽管视频生成类世界模型前景广阔,但其依赖像素级重建常引入视觉冗余,阻碍动作解码与泛化。隐空间世界模型虽具紧凑性与运动感知能力,却忽略精细操作所需的细节。为此,我们提出MoWM,一种融合异构世界模型表示的混合世界模型框架。该方法结合运动感知隐空间特征与像素空间特征,使模型能聚焦于对动作解码关键的视觉细节。在CALVIN数据集及真实世界操作任务上的广泛评估表明,本方法实现了最先进的任务成功率与优异的泛化性能。我们还系统分析了两种特征空间的优势,为未来具身规划研究提供重要参考。代码已开源:https://github.com/tsinghua-fib-lab/MoWM。

原文摘要 · Abstract (English)

Embodied action planning is a core challenge in robotics, requiring models to generate precise actions from visual observations and language instructions. While video generation world models are promising, their reliance on pixel-level reconstruction often introduces visual redundancies that hinder action decoding and generalization. Latent world models offer a compact, motion-aware representation, but overlook the fine-grained details critical for precise manipulation. To overcome these limitations, we propose MoWM, a mixture-of-world-model framework that fuses representations from hybrid world models for embodied action planning. Our approach combines motion-aware latent world model features with pixel-space features, enabling MoWM to emphasize action-relevant visual details for action decoding. Extensive evaluations on the CALVIN and real-world manipulation tasks demonstrate that our method achieves state-of-the-art task success rates and superior generalization. We also provide a comprehensive analysis of the strengths of each feature space, offering valuable insights for future research in embodied planning. The code is available at: https://github.com/tsinghua-fib-lab/MoWM.

具身规划世界模型机器人多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。