arXiv:2509.09547cs.CVcs.AI2025-09被引 1

通过融合自监督视觉编码器特征,提升视频生成模型质量。

Improving Video Diffusion Transformer Training by Multi-Feature Fusion and Alignment from Self-Supervised Vision Encoders

  • 引入多特征融合与对齐机制,增强生成模型中间特征表达
  • 在无条件与类别条件生成任务上均显著提升视频质量
  • 适合关注视频生成质量优化的研究者与开发者

视频扩散模型近年来因架构创新(如扩散变压器)和新型训练目标(如流匹配)而快速发展,但对其特征表示能力的改进仍不足。本文提出,通过将视频生成器的中间特征与预训练视觉编码器的特征进行对齐,可有效提升训练效果。我们设计了一种新评估指标,深入分析多种自监督视觉编码器的判别性与时间一致性,以判断其适配性。基于此分析,提出Align4Gen方法,将多特征融合与对齐机制融入视频扩散模型训练流程。在无条件与类别条件视频生成任务上的实验表明,该方法在多个量化指标上均取得显著提升。完整生成结果见项目主页:https://align4gen.github.io/align4gen/

原文摘要 · Abstract (English)

Video diffusion models have advanced rapidly in the recent years as a result of series of architectural innovations (e.g., diffusion transformers) and use of novel training objectives (e.g., flow matching). In contrast, less attention has been paid to improving the feature representation power of such models. In this work, we show that training video diffusion models can benefit from aligning the intermediate features of the video generator with feature representations of pre-trained vision encoders. We propose a new metric and conduct an in-depth analysis of various vision encoders to evaluate their discriminability and temporal consistency, thereby assessing their suitability for video feature alignment. Based on the analysis, we present Align4Gen which provides a novel multi-feature fusion and alignment method integrated into video diffusion model training. We evaluate Align4Gen both for unconditional and class-conditional video generation tasks and show that it results in improved video generation as quantified by various metrics. Full video results are available on our project page: https://align4gen.github.io/align4gen/

视频生成扩散模型特征对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。