让视频生成模型更平滑地实现属性渐变。
From Prompt to Progression: Taming Video Diffusion Models for Seamless Attribute Transition
- 在去噪过程中引入逐帧引导,实现属性渐变
- 新基准测试显示优于现有方法,过渡更连贯
- 适合需要精确控制视频属性变化的研究者
现有模型在生成具有渐进属性变化的视频时表现不佳,常见提示插值方法难以处理属性渐变,导致不一致问题加剧。本文提出一种简单有效的方法,在去噪过程中引入逐帧引导,为每个噪声潜在表示构建数据特定的过渡方向,实现从初始到最终属性的逐帧平滑转移,同时保持视频运动动态。我们还提出了受控属性过渡基准(CAT-Bench),融合属性与运动动态,全面评估模型性能,并设计两个指标衡量属性转换的准确性和平滑性。实验结果表明,该方法在视觉保真度、提示对齐性和属性过渡连贯性方面均优于现有基线。代码与数据集已公开:https://github.com/lynn-ling-lo/Prompt2Progression。
原文摘要 · Abstract (English)
Existing models often struggle with complex temporal changes, particularly when generating videos with gradual attribute transitions. The most common prompt interpolation approach for motion transitions often fails to handle gradual attribute transitions, where inconsistencies tend to become more pronounced. In this work, we propose a simple yet effective method to extend existing models for smooth and consistent attribute transitions, through introducing frame-wise guidance during the denoising process. Our approach constructs a data-specific transitional direction for each noisy latent, guiding the gradual shift from initial to final attributes frame by frame while preserving the motion dynamics of the video. Moreover, we present the Controlled-Attribute-Transition Benchmark (CAT-Bench), which integrates both attribute and motion dynamics, to comprehensively evaluate the performance of different models. We further propose two metrics to assess the accuracy and smoothness of attribute transitions. Experimental results demonstrate that our approach performs favorably against existing baselines, achieving visual fidelity, maintaining alignment with text prompts, and delivering seamless attribute transitions. Code and CATBench are released: https://github.com/lynn-ling-lo/Prompt2Progression.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。