用视频扩散模型实现物体运动与交互的精准控制。
Learning to Generate Rigid Body Interactions with Video Diffusion Models
- 通过物体遮罩逐步去除未来运动监督,实现物理合理运动生成。
- 在真实场景中成功模拟刚体与手物交互,性能优于同类模型。
- 支持低级运动控制与高级文本条件结合,适合机器人仿真应用。
近期视频生成模型在电影、社交媒体和广告领域取得显著进展,并展现出作为机器人与具身决策世界模拟器的潜力。然而,现有方法仍难以生成物理上合理的物体交互,且缺乏对象级控制机制。为此,我们提出KineMask,一种支持真实刚体控制、交互与效应生成的视频生成方法。给定单张图像和指定物体速度,该方法可生成包含推断运动与未来交互的视频。我们采用两阶段训练策略,通过物体掩码逐步移除未来运动监督,在合成简单交互场景上训练视频扩散模型(VDMs),并证明其在真实场景中刚体与手物交互上的显著改进与泛化能力。此外,KineMask通过预测场景描述将低级运动控制与高级文本条件结合,支持复杂动力学现象合成。实验表明,KineMask可适配多种VDMs,性能优于同等规模的最新模型。消融研究进一步揭示了低级与高级条件在VDM中的互补作用。
原文摘要 · Abstract (English)
Recent video generation models have achieved remarkable progress and are now deployed in film, social media production, and advertising. Beyond their creative potential, such models also hold promise as world simulators for robotics and embodied decision making. Despite strong advances, current approaches still struggle to generate physically plausible object interactions and lack object-level control mechanisms. To address these limitations, we introduce KineMask, an approach for video generation that enables realistic rigid body control, interactions, and effects. Given a single image and a specified object velocity, our method generates videos with inferred motions and future object interactions. We propose a two-stage training strategy that gradually removes future motion supervision via object masks. Using this strategy we train video diffusion models (VDMs) on synthetic scenes of simple interactions and demonstrate significant improvements and generalization to rigid body and hand-object interactions in real scenes. Furthermore, KineMask integrates low-level motion control with high-level textual conditioning via predicted scene descriptions, leading to support for synthesis of complex dynamical phenomena. Our experiments show that KineMask generalizes to different VDMs and achieves strong improvements over recent models of comparable size. Ablation studies further highlight the complementary roles of low- and high-level conditioning in VDMs. Project Page: https://daromog.github.io/KineMask/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。