arXiv:2501.08453cs.CVcs.LG2025-01被引 59

Vchitect-2.0用并行变压器提升长视频生成质量与训练效率

Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models

  • 采用多模态扩散块,对齐文本与视频帧,保持时间连贯性
  • 通过混合并行与内存优化,支持分布式系统上长视频高效训练
  • 构建百万级高质量数据集,适合高保真视频生成任务

我们提出Vchitect-2.0,一种专为大规模文本到视频生成设计的并行变压器架构。系统包含三项核心设计:(1) 引入新型多模态扩散块,实现文本描述与生成视频帧的一致对齐,同时保持序列间的时间连贯性;(2) 提出内存高效的训练框架,结合混合并行与多种内存压缩技术,使长视频序列可在分布式系统上高效训练;(3) 改进数据处理流程,构建Vchitect T2V DataVerse,一个经严格标注与美学评估的百万级高质量训练数据集。大量基准测试表明,Vchitect-2.0在视频质量、训练效率和可扩展性方面均优于现有方法,可作为高保真视频生成的良好基础。

原文摘要 · Abstract (English)

We present Vchitect-2.0, a parallel transformer architecture designed to scale up video diffusion models for large-scale text-to-video generation. The overall Vchitect-2.0 system has several key designs. (1) By introducing a novel Multimodal Diffusion Block, our approach achieves consistent alignment between text descriptions and generated video frames, while maintaining temporal coherence across sequences. (2) To overcome memory and computational bottlenecks, we propose a Memory-efficient Training framework that incorporates hybrid parallelism and other memory reduction techniques, enabling efficient training of long video sequences on distributed systems. (3) Additionally, our enhanced data processing pipeline ensures the creation of Vchitect T2V DataVerse, a high-quality million-scale training dataset through rigorous annotation and aesthetic evaluation. Extensive benchmarking demonstrates that Vchitect-2.0 outperforms existing methods in video quality, training efficiency, and scalability, serving as a suitable base for high-fidelity video generation.

视频生成扩散模型并行计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。