通过时序对象中心表示,实现高质量可控视频生成与编辑。
Compositional Video Synthesis by Temporal Object-Centric Learning
- 学习姿态不变的对象槽,结合预训练扩散模型生成视频。
- 在视频生成质量与时间一致性上超越现有方法。
- 支持物体插入、删除等直观编辑,适合内容创作与交互设计。
我们提出一种新的组合式视频生成框架,利用时序一致的对象中心表示,将先前的图像级方法SlotAdapt扩展至视频领域。现有对象中心方法或完全缺乏生成能力,或整体处理视频序列,忽略显式的对象级结构。我们的方法通过学习姿态不变的对象中心槽,并将其条件化于预训练扩散模型,显式捕捉时序动态。该设计实现了高质量、像素级的视频合成,具备出色的时序连贯性,并支持物体插入、删除或替换等直观的组合编辑,保持帧间物体身份一致。大量实验表明,本方法在视频生成质量和时间一致性上达到新基准,优于以往对象中心生成方法。尽管分割性能接近当前最优水平,但本方法独特地将此能力与强大的生成性能结合,显著推动了交互式与可控视频生成的发展,为高级内容创作、语义编辑及动态场景理解开辟新可能。
原文摘要 · Abstract (English)
We present a novel framework for compositional video synthesis that leverages temporally consistent object-centric representations, extending our previous work, SlotAdapt, from images to video. While existing object-centric approaches either lack generative capabilities entirely or treat video sequences holistically, thus neglecting explicit object-level structure, our approach explicitly captures temporal dynamics by learning pose invariant object-centric slots and conditioning them on pretrained diffusion models. This design enables high-quality, pixel-level video synthesis with superior temporal coherence, and offers intuitive compositional editing capabilities such as object insertion, deletion, or replacement, maintaining consistent object identities across frames. Extensive experiments demonstrate that our method sets new benchmarks in video generation quality and temporal consistency, outperforming previous object-centric generative methods. Although our segmentation performance closely matches state-of-the-art methods, our approach uniquely integrates this capability with robust generative performance, significantly advancing interactive and controllable video generation and opening new possibilities for advanced content creation, semantic editing, and dynamic scene understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。