Allegro突破视频生成质量瓶颈,逼近商业级表现
Allegro: Open the Black Box of Commercial-Level Video Generation Model
- 提出端到端训练框架,优化数据、架构与评估流程
- 用户测试显示性能仅次于Hailuo和Kling,超越多数开源模型
- 公开代码与模型,助力社区提升视频生成能力
视频生成领域虽取得显著进展,开源社区也推出了大量高质量模型和工具,但现有资源仍难以达到商业级性能。本文揭开黑箱,介绍Allegro——一款在质量和时序一致性上均表现出色的先进视频生成模型。我们系统梳理了当前领域的关键挑战,提出涵盖数据、模型架构、训练流程与评估在内的完整训练方法论。用户研究表明,Allegro性能超越现有开源模型,并接近主流商业模型,仅略逊于Hailuo与Kling。项目代码已开源(https://github.com/rhymes-ai/Allegro),模型可访问(https://huggingface.co/rhymes-ai/Allegro),官方画廊展示生成效果(https://rhymes.ai/allegro_gallery)。
原文摘要 · Abstract (English)
Significant advancements have been made in the field of video generation, with the open-source community contributing a wealth of research papers and tools for training high-quality models. However, despite these efforts, the available information and resources remain insufficient for achieving commercial-level performance. In this report, we open the black box and introduce $\textbf{Allegro}$, an advanced video generation model that excels in both quality and temporal consistency. We also highlight the current limitations in the field and present a comprehensive methodology for training high-performance, commercial-level video generation models, addressing key aspects such as data, model architecture, training pipeline, and evaluation. Our user study shows that Allegro surpasses existing open-source models and most commercial models, ranking just behind Hailuo and Kling. Code: https://github.com/rhymes-ai/Allegro , Model: https://huggingface.co/rhymes-ai/Allegro , Gallery: https://rhymes.ai/allegro_gallery .
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。