无需训练即可实现多视角复杂动作编辑,利用光流保持画面一致性。
MotionDiff: Training-free Zero-shot Interactive Motion Editing via Flow-assisted Multi-view Diffusion
- 通过光流估计和点运动模型,实现多视角动作生成
- 在不重新训练的情况下,支持旋转、拉伸等复杂动作
- 适合需要快速交互式动作修改的场景
生成模型在内容生成方面取得了显著进展,但可控编辑仍因输出固有的不确定性而困难,尤其在涉及空间信息的动作编辑中更为突出。现有基于物理的方法通常仅处理单视角简单动作(如平移、拖动),难以应对复杂旋转与拉伸动作,且难以保证多视角一致性,常需耗时重训练。为此,我们提出MotionDiff,一种无需训练、零样本的扩散方法,借助光流实现复杂多视角动作编辑。给定静态场景,用户可交互选择目标物体并添加运动先验。通过点运动模型(PKM)在多视角光流估计阶段(MFES)估计对应多视角光流,并在多视角运动扩散阶段(MMDS)利用解耦运动表示生成多视角运动结果。大量实验表明,MotionDiff在生成高质量、多视角一致的动作结果方面优于其他基于物理的生成方法。值得注意的是,该方法无需重训练,可便捷适配多种下游任务。
原文摘要 · Abstract (English)
Generative models have made remarkable advancements and are capable of producing high-quality content. However, performing controllable editing with generative models remains challenging, due to their inherent uncertainty in outputs. This challenge is praticularly pronounced in motion editing, which involves the processing of spatial information. While some physics-based generative methods have attempted to implement motion editing, they typically operate on single-view images with simple motions, such as translation and dragging. These methods struggle to handle complex rotation and stretching motions and ensure multi-view consistency, often necessitating resource-intensive retraining. To address these challenges, we propose MotionDiff, a training-free zero-shot diffusion method that leverages optical flow for complex multi-view motion editing. Specifically, given a static scene, users can interactively select objects of interest to add motion priors. The proposed Point Kinematic Model (PKM) then estimates corresponding multi-view optical flows during the Multi-view Flow Estimation Stage (MFES). Subsequently, these optical flows are utilized to generate multi-view motion results through decoupled motion representation in the Multi-view Motion Diffusion Stage (MMDS). Extensive experiments demonstrate that MotionDiff outperforms other physics-based generative motion editing methods in achieving high-quality multi-view consistent motion results. Notably, MotionDiff does not require retraining, enabling users to conveniently adapt it for various down-stream tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。