提升视频生成模型微调时的帧间语义一致性
Cross-Frame Representation Alignment for Fine-Tuning Video Diffusion Models
- 跨帧对齐隐藏状态与邻近帧特征,增强时序连贯性
- 在CogVideoX-5B等模型上显著提升视觉质量与帧间一致
- 适合需要高质量可控视频生成的研究者使用
用户级微调视频扩散模型以生成反映训练数据特定属性的视频面临显著挑战,但相关研究仍不充分。现有方法如表示对齐(REPA)虽能提升收敛速度与图像质量,但在视频生成中难以保持帧间语义一致性。为此,本文提出跨帧表示对齐(CREPA),通过将当前帧的隐藏状态与相邻帧的外部预训练视觉特征对齐,实现时序一致性增强。在CogVideoX-5B和Hunyuan Video等大规模视频扩散模型上的实验表明,结合LoRA等参数高效微调方法,CREPA显著提升了视觉保真度与跨帧语义连贯性。该方法在多种具有不同属性的数据集上均表现良好,验证了其广泛适用性。
原文摘要 · Abstract (English)
Fine-tuning Video Diffusion Models (VDMs) at the user level to generate videos that reflect specific attributes of training data presents notable challenges, yet remains underexplored despite its practical importance. Meanwhile, recent work such as Representation Alignment (REPA) has shown promise in improving the convergence and quality of DiT-based image diffusion models by aligning, or assimilating, its internal hidden states with external pretrained visual features, suggesting its potential for VDM fine-tuning. In this work, we first propose a straightforward adaptation of REPA for VDMs and empirically show that, while effective for convergence, it is suboptimal in preserving semantic consistency across frames. To address this limitation, we introduce Cross-frame Representation Alignment (CREPA), a novel regularization technique that aligns hidden states of a frame with external features from neighboring frames. Empirical evaluations on large-scale VDMs, including CogVideoX-5B and Hunyuan Video, demonstrate that CREPA improves both visual fidelity and cross-frame semantic coherence when fine-tuned with parameter-efficient methods such as LoRA. We further validate CREPA across diverse datasets with varying attributes, confirming its broad applicability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。