arXiv:2412.09551cs.CV2024-12被引 7

用示范视频生成符合物理规律的连贯新视频。

Video Creation by Demonstration

  • 通过预测未来帧实现无监督训练,用隐式动作控制生成视频。
  • 在人类偏好和机器评估中均优于现有基线模型。
  • 适合交互式世界模拟、视频创作等需要灵活控制的场景。

我们探索了一种新颖的视频生成体验——视频示范生成(Video Creation by Demonstration)。给定一段示范视频和来自不同场景的上下文图像,模型生成一个在语义上自然延续上下文图像并执行示范动作概念的物理合理视频。为实现该能力,我们提出δ-Diffusion,一种基于条件未来帧预测的自监督训练方法,利用未标注视频学习。与依赖显式信号的现有视频生成控制不同,本方法采用隐式潜在控制,以满足通用视频所需的灵活性与表现力。通过在视频基础模型上引入外观瓶颈设计,从示范视频中提取动作潜在表示,实现生成过程中的最小外观泄露。实验表明,δ-Diffusion在人类偏好和大规模机器评估中均优于相关基线,并展现出交互式世界模拟的潜力。示例生成结果可访问 https://delta-diffusion.github.io/。

原文摘要 · Abstract (English)

We explore a novel video creation experience, namely Video Creation by Demonstration. Given a demonstration video and a context image from a different scene, we generate a physically plausible video that continues naturally from the context image and carries out the action concepts from the demonstration. To enable this capability, we present $δ$-Diffusion, a self-supervised training approach that learns from unlabeled videos by conditional future frame prediction. Unlike most existing video generation controls that are based on explicit signals, we adopts the form of implicit latent control for maximal flexibility and expressiveness required by general videos. By leveraging a video foundation model with an appearance bottleneck design on top, we extract action latents from demonstration videos for conditioning the generation process with minimal appearance leakage. Empirically, $δ$-Diffusion outperforms related baselines in terms of both human preference and large-scale machine evaluations, and demonstrates potentials towards interactive world simulation. Sampled video generation results are available at https://delta-diffusion.github.io/.

视频生成隐式控制自监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。