通过融合多层特征提升视频生成的时空一致性
RepVideo: Rethinking Cross-Layer Representation for Video Generation
- 融合相邻层特征构建更稳定的语义表示
- 在多个数据集上显著提升帧间连贯性与空间细节
- 适合关注视频生成质量与时序一致性的研究者
视频生成得益于扩散模型的引入,已取得显著进展,但当前研究多聚焦于模型规模扩展,对表征本身如何影响生成过程关注不足。本文分析中间层特征发现,不同层注意力图差异显著,导致语义表示不稳定,并引发特征累积偏差,降低相邻帧相似度,影响时间连贯性。为此,我们提出RepVideo,一种增强文本到视频扩散模型的表示框架。通过聚合邻近层特征形成丰富表示,提升注意力机制输入的语义表达力,同时保障相邻帧特征一致性。大量实验表明,RepVideo不仅显著增强对复杂物体间空间关系等细节的刻画能力,还有效改善视频生成的时间一致性。
原文摘要 · Abstract (English)
Video generation has achieved remarkable progress with the introduction of diffusion models, which have significantly improved the quality of generated videos. However, recent research has primarily focused on scaling up model training, while offering limited insights into the direct impact of representations on the video generation process. In this paper, we initially investigate the characteristics of features in intermediate layers, finding substantial variations in attention maps across different layers. These variations lead to unstable semantic representations and contribute to cumulative differences between features, which ultimately reduce the similarity between adjacent frames and negatively affect temporal coherence. To address this, we propose RepVideo, an enhanced representation framework for text-to-video diffusion models. By accumulating features from neighboring layers to form enriched representations, this approach captures more stable semantic information. These enhanced representations are then used as inputs to the attention mechanism, thereby improving semantic expressiveness while ensuring feature consistency across adjacent frames. Extensive experiments demonstrate that our RepVideo not only significantly enhances the ability to generate accurate spatial appearances, such as capturing complex spatial relationships between multiple objects, but also improves temporal consistency in video generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。