超大规模视频生成模型,专为营销场景打造
Aquarius: A Family of Industry-Level Video Generation Models for Marketing Scenarios
- 基于分布式架构与高效算法,支持百亿参数级视频生成
- 长时长、多画幅、高保真视频合成,训练效率达36% MFU
- 适合广告、电商等工业级视频生成需求,开源数据处理框架
本文介绍Aquarius,一套面向营销场景的工业级视频生成模型家族,专为数千xPU集群和数百亿参数模型设计。通过高效的工程架构与算法创新,Aquarius在高保真、多画幅比例和长时长视频合成方面表现卓越。其框架包含五大组件:分布式图与视频数据处理流水线,可自动调度数万CPU与数千xPU,实现高效视频数据处理;后续将开源名为“Aquarius-Datapipe”的完整数据处理框架。不同规模的模型架构包括2B参数的Single-DiT和13.4B参数的Multimodal-DiT,支持多画幅、多分辨率、多时长视频生成。高性能训练基础设施采用混合并行与细粒度内存优化策略,在大规模下达到36% MFU。多xPU并行推理加速通过扩散缓存与注意力优化,实现2.35倍推理速度提升。已应用于图像转视频、文本转视频(虚拟人)、视频修复与个性化等营销场景,后续将扩展更多下游应用与多维评估指标。
原文摘要 · Abstract (English)
This report introduces Aquarius, a family of industry-level video generation models for marketing scenarios designed for thousands-xPU clusters and models with hundreds of billions of parameters. Leveraging efficient engineering architecture and algorithmic innovation, Aquarius demonstrates exceptional performance in high-fidelity, multi-aspect-ratio, and long-duration video synthesis. By disclosing the framework's design details, we aim to demystify industrial-scale video generation systems and catalyze advancements in the generative video community. The Aquarius framework consists of five components: Distributed Graph and Video Data Processing Pipeline: Manages tens of thousands of CPUs and thousands of xPUs via automated task distribution, enabling efficient video data processing. Additionally, we are about to open-source the entire data processing framework named "Aquarius-Datapipe". Model Architectures for Different Scales: Include a Single-DiT architecture for 2B models and a Multimodal-DiT architecture for 13.4B models, supporting multi-aspect ratios, multi-resolution, and multi-duration video generation. High-Performance infrastructure designed for video generation model training: Incorporating hybrid parallelism and fine-grained memory optimization strategies, this infrastructure achieves 36% MFU at large scale. Multi-xPU Parallel Inference Acceleration: Utilizes diffusion cache and attention optimization to achieve a 2.35x inference speedup. Multiple marketing-scenarios applications: Including image-to-video, text-to-video (avatar), video inpainting and video personalization, among others. More downstream applications and multi-dimensional evaluation metrics will be added in the upcoming version updates.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。