用扭曲噪声训练视频扩散模型,让生成更连贯且采样更快。
On Equivariance and Fast Sampling in Video Diffusion Models Trained with Warped Noise
- 用扭曲噪声配合标准去噪目标,自动实现输入变换的等变性。
- 采样步数减少一半仍保持高画质,生成运动更连贯。
- 适合需要高效、精准控制视频运动的应用场景。
时间一致的视频到视频生成对风格迁移和超分辨率等应用至关重要。本文对最近提出的扭曲噪声训练方法进行了理论分析,发现将其与标准去噪目标结合,可隐式训练出对输入噪声空间变换具备等变性的模型,称为EquiVDM。这种等变性使输入噪声中的运动自然对齐生成视频中的运动,无需额外模块或辅助损失即可获得连贯、高质量输出。另一个优势是采样效率:EquiVDM在远少于传统方法的采样步数下达到相当或更优质量。当压缩为单步学生模型时,EquiVDM保持等变性,相比非等变基线展现出更强的运动可控性和保真度。在多个基准上,EquiVDM在运动对齐、时间一致性与感知质量方面持续优于先前方法,同时显著降低采样成本。
原文摘要 · Abstract (English)
Temporally consistent video-to-video generation is critical for applications such as style transfer and upsampling. In this paper, we provide a theoretical analysis of warped noise - a recently proposed technique for training video diffusion models - and show that pairing it with the standard denoising objective implicitly trains models to be equivariant to spatial transformations of the input noise, which we term EquiVDM. This equivariance enables motion in the input noise to align naturally with motion in the generated video, yielding coherent, high-fidelity outputs without the need for specialized modules or auxiliary losses. A further advantage is sampling efficiency: EquiVDM achieves comparable or superior quality in far fewer sampling steps. When distilled into one-step student models, EquiVDM preserves equivariance and delivers stronger motion controllability and fidelity than distilled nonequivariant baselines. Across benchmarks, EquiVDM consistently outperforms prior methods in motion alignment, temporal consistency, and perceptual quality, while substantially lowering sampling cost.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。