arXiv:2412.00148cs.CV2024-12CVPR被引 14

无需训练,从静态图生成多样真实物体运动

Motion Modes: What Could Happen Next?

  • 用能量函数引导流生成器分离物体与镜头运动
  • 不依赖训练数据,生成动作多样性超越现有方法
  • 适合需要多样化视频生成的场景应用

从单张静态图像预测多样化的物体运动仍具挑战性,因现有视频生成模型常将物体运动与摄像机运动及其他场景变化混淆。尽管近期方法能基于运动箭头输入生成特定运动,但依赖合成数据和预定义动作,限制了在复杂场景中的应用。我们提出Motion Modes,一种无需训练的方法,通过探索预训练图像到视频生成器的潜在分布,发现静态图像中选定物体的多种独特且合理运动。该方法利用由能量函数引导的流生成器,实现物体与相机运动的解耦;同时采用受粒子引导启发的能量机制,进一步提升生成动作的多样性,无需显式训练数据。实验表明,Motion Modes生成的动画既真实又多样,在可合理性与多样性上优于先前方法,甚至超过人类判断。

原文摘要 · Abstract (English)

Predicting diverse object motions from a single static image remains challenging, as current video generation models often entangle object movement with camera motion and other scene changes. While recent methods can predict specific motions from motion arrow input, they rely on synthetic data and predefined motions, limiting their application to complex scenes. We introduce Motion Modes, a training-free approach that explores a pre-trained image-to-video generator's latent distribution to discover various distinct and plausible motions focused on selected objects in static images. We achieve this by employing a flow generator guided by energy functions designed to disentangle object and camera motion. Additionally, we use an energy inspired by particle guidance to diversify the generated motions, without requiring explicit training data. Experimental results demonstrate that Motion Modes generates realistic and varied object animations, surpassing previous methods and even human predictions regarding plausibility and diversity. Project Webpage: https://motionmodes.github.io/

视频生成运动预测无训练多样性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。