让视频生成更真实,机器人操作更可靠。
GEM-4D: Geometry-Enhanced Video World Models for Robot Manipulation

- 用几何监督让视频保持物体位置一致
- 真实世界操作成功率从61%提升到81%
- 适合需要精准动作的机器人研究者
视频世界模型能从单一指令生成未来画面,但常无法在时间上持续追踪同一物理点。为此,我们提出GEM-4D,一种几何增强型视频世界模型,在训练时注入来自预训练几何基础模型的密集4D对应关系监督,使模型在单流架构下同时捕捉外观与几何结构,无额外推理开销。我们还引入逆动力学模块,将一致性视频推演转化为可执行的机器人轨迹,支持在真实与仿真环境中直接部署。GEM-4D在视频预测与几何一致性上均达到当前最优表现,真实世界操作成功率由61%提升至81%。更多结果见https://gem-4d.github.io/。
原文摘要 · Abstract (English)
Video world models can generate realistic futures from a single instruction, but they often fail to track the same physical points consistently across time. As a result, the generated videos appear plausible, yet lack the physical grounding required for reliable action execution, such as robot manipulation. We present GEM-4D, a geometry-grounded video world model that resolves this limitation by injecting dense 4D correspondence supervision distilled from a pretrained geometry foundation model into the video generative backbone during training. This supervision enables the model to jointly capture appearance and geometric structure while retaining a single-stream architecture with no additional inference cost. We further introduce an inverse dynamics module that converts correspondence-consistent video rollouts into executable robot trajectories, enabling direct deployment in both real-world and simulated manipulation. GEM-4D achieves state-of-the-art performance on both video prediction and geometric consistency across both simulation and realistic scenarios and improves real-world manipulation success from 61% to 81%. Additional results are available at https://gem-4d.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。