开源大模型可生成高分辨率长视频,支持多种输入条件。
Open-Sora Plan: Open-Source Large Video Generation Model
- 采用小波流变分自编码器与图像视频联合去噪器,提升生成质量。
- 支持多条件控制,实现高质量长时高分辨率视频生成。
- 提供完整代码与权重,适合研究者快速复现和二次开发。
我们提出 Open-Sora Plan,一个开源项目,旨在构建一个大型视频生成模型,可根据多种用户输入生成高分辨率、长时长视频。该系统包含完整的视频生成流程组件:小波流变分自编码器(Wavelet-Flow Variational Autoencoder)、图像-视频联合稀疏去噪器(Joint Image-Video Sparse Denoiser)以及多种条件控制器。同时,设计了高效的训练与推理辅助策略,并提出多维度数据清洗管道以获取高质量数据。得益于高效的设计思路,Open-Sora Plan 在定性与定量评估中均表现出色。所有代码与模型权重已公开于 https://github.com/PKU-YuanGroup/Open-Sora-Plan,期待为视频生成研究社区提供启发。
原文摘要 · Abstract (English)
We introduce Open-Sora Plan, an open-source project that aims to contribute a large generation model for generating desired high-resolution videos with long durations based on various user inputs. Our project comprises multiple components for the entire video generation process, including a Wavelet-Flow Variational Autoencoder, a Joint Image-Video Skiparse Denoiser, and various condition controllers. Moreover, many assistant strategies for efficient training and inference are designed, and a multi-dimensional data curation pipeline is proposed for obtaining desired high-quality data. Benefiting from efficient thoughts, our Open-Sora Plan achieves impressive video generation results in both qualitative and quantitative evaluations. We hope our careful design and practical experience can inspire the video generation research community. All our codes and model weights are publicly available at \url{https://github.com/PKU-YuanGroup/Open-Sora-Plan}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。