让视频生成能根据边界条件推理物理合理运动
Motion Dreamer: Boundary Conditional Motion Reasoning for Physically Coherent Video Generation
- 分两阶段设计,先推理运动再生成画面
- 用稀疏到稠密的运动表示,支持部分输入运动
- 适合自动驾驶等需精准运动预测的场景
视频生成在自动驾驶和具身智能的规划控制中具有重要意义。然而真实应用不仅要求视觉合理,还需基于显式边界条件(如初始场景图与部分物体运动)进行运动推理。现有方法或忽略用户定义的运动约束,导致物理不一致;或要求完整运动输入,实际难以获取。本文提出 Motion Dreamer,一种两阶段框架,将运动推理与视觉合成分离。引入实例流(instance flow)这一稀疏到稠密的运动表示,有效融合部分用户定义的运动,并采用运动补全策略,实现对其他物体运动的鲁棒推理。大量实验表明,Motion Dreamer 显著优于现有方法,在运动合理性与视觉真实感上均表现更优,推动了实用化边界条件运动推理的发展。
原文摘要 · Abstract (English)
Recent advances in video generation have shown promise for generating future scenarios, critical for planning and control in autonomous driving and embodied intelligence. However, real-world applications demand more than visually plausible predictions; they require reasoning about object motions based on explicitly defined boundary conditions, such as initial scene image and partial object motion. We term this capability Boundary Conditional Motion Reasoning. Current approaches either neglect explicit user-defined motion constraints, producing physically inconsistent motions, or conversely demand complete motion inputs, which are rarely available in practice. Here we introduce Motion Dreamer, a two-stage framework that explicitly separates motion reasoning from visual synthesis, addressing these limitations. Our approach introduces instance flow, a sparse-to-dense motion representation enabling effective integration of partial user-defined motions, and the motion inpainting strategy to robustly enable reasoning motions of other objects. Extensive experiments demonstrate that Motion Dreamer significantly outperforms existing methods, achieving superior motion plausibility and visual realism, thus bridging the gap towards practical boundary conditional motion reasoning. Our webpage is available: https://envision-research.github.io/MotionDreamer/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。