开源视频生成模型 HunyuanVideo 性能媲美甚至超越闭源顶尖模型。
HunyuanVideo: A Systematic Framework For Large Video Generative Models
- 构建端到端开源框架,涵盖数据、架构、训练与推理优化。
- 训练出超130亿参数模型,是当前最大的开源视频生成模型。
- 适合研究者与开发者使用,推动视频生成生态开放创新。
近期视频生成技术进步深刻影响了个人与产业应用,但主流模型仍为闭源,导致公众可用能力与产业水平存在显著差距。本文介绍 HunyuanVideo,一个创新的开源视频基础模型,其生成性能可媲美甚至超越领先闭源模型。该模型包含完整框架,整合数据筛选、先进架构设计、渐进式模型扩展与训练、以及面向大规模训练与推理的高效基础设施。我们成功训练出参数量超过130亿的视频生成模型,成为目前最大的开源视频生成模型。通过大量实验与针对性设计,保障了高视觉质量、流畅运动、文图对齐及高级影视手法表现。专业评估显示,HunyuanVideo 在性能上优于 Runway Gen-3、Luma 1.6 及三款顶尖中文视频生成模型。项目代码已公开于 https://github.com/Tencent/HunyuanVideo,旨在弥合闭源与开源社区差距,赋能开发者探索创意,共建活跃的视频生成生态。
原文摘要 · Abstract (English)
Recent advancements in video generation have significantly impacted daily life for both individuals and industries. However, the leading video generation models remain closed-source, resulting in a notable performance gap between industry capabilities and those available to the public. In this report, we introduce HunyuanVideo, an innovative open-source video foundation model that demonstrates performance in video generation comparable to, or even surpassing, that of leading closed-source models. HunyuanVideo encompasses a comprehensive framework that integrates several key elements, including data curation, advanced architectural design, progressive model scaling and training, and an efficient infrastructure tailored for large-scale model training and inference. As a result, we successfully trained a video generative model with over 13 billion parameters, making it the largest among all open-source models. We conducted extensive experiments and implemented a series of targeted designs to ensure high visual quality, motion dynamics, text-video alignment, and advanced filming techniques. According to evaluations by professionals, HunyuanVideo outperforms previous state-of-the-art models, including Runway Gen-3, Luma 1.6, and three top-performing Chinese video generative models. By releasing the code for the foundation model and its applications, we aim to bridge the gap between closed-source and open-source communities. This initiative will empower individuals within the community to experiment with their ideas, fostering a more dynamic and vibrant video generation ecosystem. The code is publicly available at https://github.com/Tencent/HunyuanVideo.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。