用新指标提升图像生成视频的连贯性,不降质还更稳。
Video Consistency Distance: Enhancing Temporal Consistency for Image-to-Video Generation via Reward-Based Fine-Tuning
- 在频域分析帧特征,设计新度量增强时间连贯性
- 多数据集实验表明连贯性显著提升,其他质量不变
- 适合关注视频生成稳定性的研究者和开发者
基于奖励的微调是提升视频扩散模型生成质量的有效方法,无需真实视频数据集即可优化模型。然而,传统奖励函数主要关注整体视频质量(如美学和整体一致性),在图像到视频(I2V)生成任务中往往导致时间连贯性下降。为此,我们提出视频一致性距离(Video Consistency Distance, VCD),一种专为增强时间连贯性设计的新度量,并在基于奖励的微调框架下进行模型优化。VCD在视频帧特征的频域空间定义,通过频域分析有效捕捉帧间信息。多个I2V数据集的实验结果表明,使用VCD微调模型能显著提升时间连贯性,同时不损害其他性能指标。
原文摘要 · Abstract (English)
Reward-based fine-tuning of video diffusion models is an effective approach to improve the quality of generated videos, as it can fine-tune models without requiring real-world video datasets. However, it can sometimes be limited to specific performances because conventional reward functions are mainly aimed at enhancing the quality across the whole generated video sequence, such as aesthetic appeal and overall consistency. Notably, the temporal consistency of the generated video often suffers when applying previous approaches to image-to-video (I2V) generation tasks. To address this limitation, we propose Video Consistency Distance (VCD), a novel metric designed to enhance temporal consistency, and fine-tune a model with the reward-based fine-tuning framework. To achieve coherent temporal consistency relative to a conditioning image, VCD is defined in the frequency space of video frame features to capture frame information effectively through frequency-domain analysis. Experimental results across multiple I2V datasets demonstrate that fine-tuning a video generation model with VCD significantly enhances temporal consistency without degrading other performance compared to the previous method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。