开源视频生成模型Open-Sora让普通人也能高效制作高质量视频。
Open-Sora: Democratizing Efficient Video Production for All
- 采用时空分离的扩散变换器架构,提升视频生成效率。
- 可生成最长15秒、720p分辨率、任意画幅的高清视频。
- 全开源代码与模型权重,适合开发者和创意者使用。
视觉与语言是人类认知的基础能力,尽管人工智能在语言方面取得显著进展,但视觉智能——尤其是生成和模拟真实世界影像的能力——仍严重滞后。为推动人工视觉智能的发展与普及,我们推出开源视频生成模型Open-Sora,支持文本到图像、文本到视频及图像到视频等多种生成任务。该模型采用先进的深度学习架构与训练推理技术,可生成最长15秒、最高720p分辨率、任意宽高比的高保真视频内容。我们提出时空分离扩散变换器(STDiT),通过解耦空间与时间注意力提升效率;并设计高度压缩的3D自编码器,结合定制化训练策略,显著加速训练过程。本项目全面开放训练/推理/数据准备代码及模型权重,所有资源均公开于https://github.com/hpcaitech/Open-Sora,旨在促进AI内容创作领域的创新、创造力与包容性。
原文摘要 · Abstract (English)
Vision and language are the two foundational senses for humans, and they build up our cognitive ability and intelligence. While significant breakthroughs have been made in AI language ability, artificial visual intelligence, especially the ability to generate and simulate the world we see, is far lagging behind. To facilitate the development and accessibility of artificial visual intelligence, we created Open-Sora, an open-source video generation model designed to produce high-fidelity video content. Open-Sora supports a wide spectrum of visual generation tasks, including text-to-image generation, text-to-video generation, and image-to-video generation. The model leverages advanced deep learning architectures and training/inference techniques to enable flexible video synthesis, which could generate video content of up to 15 seconds, up to 720p resolution, and arbitrary aspect ratios. Specifically, we introduce Spatial-Temporal Diffusion Transformer (STDiT), an efficient diffusion framework for videos that decouples spatial and temporal attention. We also introduce a highly compressive 3D autoencoder to make representations compact and further accelerate training with an ad hoc training strategy. Through this initiative, we aim to foster innovation, creativity, and inclusivity within the community of AI content creation. By embracing the open-source principle, Open-Sora democratizes full access to all the training/inference/data preparation codes as well as model weights. All resources are publicly available at: https://github.com/hpcaitech/Open-Sora.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。