arXiv:2606.32028cs.RO2026-06被引 1

拆解视频生成与动态建模,提升机器人操作的世界模型效率

DVG-WM: Disentangled Video Generation Enables Efficient Embodied World Model for Robotic Manipulation

论文配图:DVG-WM: Disentangled Video Generation Enables Efficient Embodied World Model for Robotic Manipulation
图 1 · 摘自论文原文
  • 分离动态学习与视觉合成,分步生成交互预测
  • 实测视频生成速度提升3.97倍,细节保留更优
  • 适合需要快速推理的机器人操控场景

基于视频的具身世界模型通过预测未来状态为机器人操作提供有力支持,但现有方法受限于本质纠缠:精确建模动态需低层次时序推理,而高分辨率图像生成又依赖高层次语义的广泛视觉合成。这导致迭代规划时推理缓慢或预测过于粗糙,难以保留接触细节。为此,我们提出解耦视频生成世界模型(DVG-WM),显式将世界建模分解为动态学习与视觉合成两部分。在初始观测和语言指令条件下,模型先生成中间视觉状态序列以预览物理交互,再精细化生成高保真视频。此外,提出高效级联机制:利用流匹配直接将动态映射至视频隐空间,并引入隐空间退化机制以重建接触丰富细节。在LIBERO及真实平台上的实验表明,该模型在视频质量提升的同时实现最高3.97倍加速,验证了解耦视频生成可作为高效具身世界模型用于机器人操控。

原文摘要 · Abstract (English)

Video-based embodied world models provide an appealing substrate for robotic manipulation by predicting future states, yet current approaches remain limited by a fundamental entanglement: accurately modeling dynamics typically requires low-level temporal reasoning, while producing high-resolution frames demands expansive visual synthesis according to high-level semantics. This entanglement results in slow inference speed for iterative planning or too coarse predictions to retain contact-rich details. To solve this dilemma, we present Disentangled Video Generation World Model (DVG-WM), an efficient framework that explicitly decomposes world modeling into dynamics learning and visual synthesis. Conditioned on an initial observation and a language instruction, our model first generates a plausible sequence of intermediate visual states to preview the physical interaction and refines them to obtain high-fidelity videos. Furthermore, an efficient cascading mechanism is proposed, where DVG-WM uses flow matching to directly map the dynamics to video latents, and introduces a latent degradation mechanism to regenerate contact-rich details. Experiments on LIBERO and real-world platforms demonstrate improved video quality with up to 3.97 times acceleration, validating that disentangled video generation can be an efficient embodied world model for robotic manipulation.

世界模型视频生成机器人操控解耦建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。