arXiv:2410.04171cs.CVcs.AI2024-10ICLR被引 5

用图像模型提升视频生成质量,让帧更清晰且时序连贯。

IV-Mixed Sampler: Leveraging Image Diffusion Models for Enhanced Video Synthesis

  • 不需训练,用图像扩散模型优化每帧,视频模型保持时间一致性。
  • 在4个基准上达到顶尖性能,最高使视频质量指标降低近50点。
  • 适合想提升开源视频生成效果的研究者和开发者。

多步采样机制是视觉扩散模型的关键特征,具有显著潜力通过增加推理计算成本来提升生成性能,类似OpenAI的Strawberry。已有充分研究表明,合理扩展采样过程中的计算量可有效提升生成质量、编辑能力与组合泛化性。尽管图像生成领域已发展出大量高计算量算法,但针对视频扩散模型(VDMs)的推理扩展规律研究仍较少,且现有成果带来的性能提升几乎无法被人眼察觉。为此,我们提出一种无需训练的新算法IV-Mixed Sampler,利用图像扩散模型(IDMs)的优势辅助视频模型突破当前能力瓶颈。该方法的核心在于:由图像模型显著提升每一帧质量,而视频模型则在采样过程中保障视频的时间连贯性。实验表明,IV-Mixed Sampler在包括UCF-101-FVD、MSR-VTT-FVD、Chronomagic-Bench-150和Chronomagic-Bench-1649在内的4个基准上均达到最先进水平。例如,使用IV-Mixed Sampler的开源Animatediff将UMT-FVD得分从275.2降至228.6,逼近闭源模型Pika-2.0的223.1。

原文摘要 · Abstract (English)

The multi-step sampling mechanism, a key feature of visual diffusion models, has significant potential to replicate the success of OpenAI's Strawberry in enhancing performance by increasing the inference computational cost. Sufficient prior studies have demonstrated that correctly scaling up computation in the sampling process can successfully lead to improved generation quality, enhanced image editing, and compositional generalization. While there have been rapid advancements in developing inference-heavy algorithms for improved image generation, relatively little work has explored inference scaling laws in video diffusion models (VDMs). Furthermore, existing research shows only minimal performance gains that are perceptible to the naked eye. To address this, we design a novel training-free algorithm IV-Mixed Sampler that leverages the strengths of image diffusion models (IDMs) to assist VDMs surpass their current capabilities. The core of IV-Mixed Sampler is to use IDMs to significantly enhance the quality of each video frame and VDMs ensure the temporal coherence of the video during the sampling process. Our experiments have demonstrated that IV-Mixed Sampler achieves state-of-the-art performance on 4 benchmarks including UCF-101-FVD, MSR-VTT-FVD, Chronomagic-Bench-150, and Chronomagic-Bench-1649. For example, the open-source Animatediff with IV-Mixed Sampler reduces the UMT-FVD score from 275.2 to 228.6, closing to 223.1 from the closed-source Pika-2.0.

视频生成扩散模型图像增强推理优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。