arXiv:2503.20314cs.CV2025-03被引 2.4k

Wan是开源的大型视频生成模型,支持多任务且可在消费级显卡运行。

Wan: Open and Advanced Large-Scale Video Generative Models

  • 基于扩散Transformer架构,融合新型VAE与大规模数据训练策略。
  • 14B模型在多基准测试中超越开源与商业模型,1.3B模型仅需8.19GB显存。
  • 支持图像转视频等八类任务,全系列开源,适合研究与创意应用。

本报告介绍Wan,一个全面且开源的视频基础模型套件,旨在推动视频生成技术边界。基于主流扩散Transformer范式,Wan通过新型VAE、可扩展预训练策略、大规模数据筛选及自动化评估指标等多项创新,在生成能力上实现显著提升。具体具有四大特点:领先性能——14B参数模型在包含数十亿图像和视频的大规模数据集上训练,展现出视频生成的数据与模型规模扩展规律,各项内部外部基准测试均显著优于现有开源模型及顶尖商业方案;全面性——提供1.3B与14B两个版本,兼顾效率与效果,覆盖图像转视频、指令引导视频编辑、个人视频生成等多达八项下游任务;消费级高效——1.3B模型仅需8.19GB VRAM,兼容广泛消费级GPU;开放性——完整开源代码与所有模型,网址为https://github.com/Wan-Video/Wan2.1,致力于推动视频生成社区发展,拓展产业创作可能,并为学术界提供高质量视频基础模型。

原文摘要 · Abstract (English)

This report presents Wan, a comprehensive and open suite of video foundation models designed to push the boundaries of video generation. Built upon the mainstream diffusion transformer paradigm, Wan achieves significant advancements in generative capabilities through a series of innovations, including our novel VAE, scalable pre-training strategies, large-scale data curation, and automated evaluation metrics. These contributions collectively enhance the model's performance and versatility. Specifically, Wan is characterized by four key features: Leading Performance: The 14B model of Wan, trained on a vast dataset comprising billions of images and videos, demonstrates the scaling laws of video generation with respect to both data and model size. It consistently outperforms the existing open-source models as well as state-of-the-art commercial solutions across multiple internal and external benchmarks, demonstrating a clear and significant performance superiority. Comprehensiveness: Wan offers two capable models, i.e., 1.3B and 14B parameters, for efficiency and effectiveness respectively. It also covers multiple downstream applications, including image-to-video, instruction-guided video editing, and personal video generation, encompassing up to eight tasks. Consumer-Grade Efficiency: The 1.3B model demonstrates exceptional resource efficiency, requiring only 8.19 GB VRAM, making it compatible with a wide range of consumer-grade GPUs. Openness: We open-source the entire series of Wan, including source code and all models, with the goal of fostering the growth of the video generation community. This openness seeks to significantly expand the creative possibilities of video production in the industry and provide academia with high-quality video foundation models. All the code and models are available at https://github.com/Wan-Video/Wan2.1.

视频生成扩散模型开源轻量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。