ALIVE让视频生成支持音画同步与参考动画,效果媲美商用模型。
ALIVE: Animate Your World with Lifelike Audio-Video Generation
- 基于T2V模型扩展音视频联合生成和参考驱动动画能力。
- 在百万级高质量数据上训练后,性能超越开源模型并接近商用水平。
- 提供完整数据流程与新基准,助力社区高效开发音视频模型。
视频生成正朝着统一的音视频生成方向发展。本文提出ALIVE,一个将预训练文本到视频(T2V)模型适配为类Sora风格音视频生成与动画的模型。该模型相比基础T2V模型,新增文本到音视频(T2VA)和参考到音视频(动画)生成能力。为实现音视频同步与参考动画,我们在主流MMDiT架构中引入联合音视频分支,包含用于时序对齐跨模态融合的TA-CrossAttn,以及实现精确音视频对齐的UniTemp-RoPE。同时,设计了包含音视频描述生成、质量控制等环节的完整数据流水线,以收集高质量微调数据。此外,我们构建了一个新基准用于全面评估与对比模型。在百万级高质量数据上持续预训练和微调后,ALIVE表现出色,持续优于开源模型,并达到或超越当前最先进的商业解决方案。通过提供详细训练方案与基准,我们希望ALIVE能帮助社区更高效地推进音视频生成模型的发展。官方页面:https://github.com/FoundationVision/Alive。
原文摘要 · Abstract (English)
Video generation is rapidly evolving towards unified audio-video generation. In this paper, we present ALIVE, a generation model that adapts a pretrained Text-to-Video (T2V) model to Sora-style audio-video generation and animation. In particular, the model unlocks the Text-to-Video&Audio (T2VA) and Reference-to-Video&Audio (animation) capabilities compared to the T2V foundation models. To support the audio-visual synchronization and reference animation, we augment the popular MMDiT architecture with a joint audio-video branch which includes TA-CrossAttn for temporally-aligned cross-modal fusion and UniTemp-RoPE for precise audio-visual alignment. Meanwhile, a comprehensive data pipeline consisting of audio-video captioning, quality control, etc., is carefully designed to collect high-quality finetuning data. Additionally, we introduce a new benchmark to perform a comprehensive model test and comparison. After continue pretraining and finetuning on million-level high-quality data, ALIVE demonstrates outstanding performance, consistently outperforming open-source models and matching or surpassing state-of-the-art commercial solutions. With detailed recipes and benchmarks, we hope ALIVE helps the community develop audio-video generation models more efficiently. Official page: https://github.com/FoundationVision/Alive.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。