arXiv:2601.04151cs.CVcs.AI2026-01被引 6

Apollo统一生成音视频,解决同步差、对齐弱难题。

Apollo: Unified Multi-Task Audio-Video Joint Generation

  • 单塔架构+全注意力机制,实现音视频紧密对齐。
  • 多任务渐进训练避免单模态退化,提升泛化能力。
  • 自动生成百万级高质量音视频三元组数据集。

音视频联合生成进展迅速,但仍存在非商业方法中音画不同步、口型与语音对不齐、单模态质量下降等问题,根源在于音频视觉对应建模弱、泛化能力差及高质量密集标注数据稀缺。为此,我们提出Apollo,从模型架构、训练策略和数据构建三方面入手:采用统一的单塔设计与全注意力机制,实现紧致音视频对齐;通过随机模态掩码与多阶段课程学习,促进多任务联合优化,增强音视频对齐的世界知识并防止单模态崩溃;构建首个大规模带密集标注的音视频数据集,并提出自动化数据构造流水线,标注筛选出数百万条高质量、严格对齐的音视频-文本三元组。基于此,Apollo在联合与单模态生成中均实现高保真、语义与时间对齐的指令跟随生成,且在分布外场景下仍表现稳健,跨任务显著优于以往方法,性能媲美Veo 3,为下一代音视频合成提供统一可扩展路径。

原文摘要 · Abstract (English)

Audio-video joint generation has progressed rapidly, yet substantial challenges still remain. Non-commercial approaches still suffer audio-visual asynchrony, poor lip-speech alignment, and unimodal degradation, which can be stemmed from weak audio-visual correspondence modeling, limited generalization, and scarce high-quality dense-caption data. To address these issues, we introduce Apollo and delve into three axes--model architecture, training strategy, and data curation. Architecturally, we adopt a single-tower design with unified DiT blocks and an Omni-Full Attention mechanism, achieving tight audio-visual alignment and strong scalability. Training-wise, we adopt a progressive multitask regime--random modality masking to joint optimization across tasks, and a multistage curriculum, yielding robust representations, strengthening A-V aligned world knowledge, and preventing unimodal collapse. For datasets, we present the first large-scale audio-video dataset with dense captions, and introduce a novel automated data-construction pipeline which annotates and filters millions of diverse, high-quality, strictly aligned audio-video-caption triplets. Building on this, Apollo scales to large datasets, delivering high-fidelity, semantically and temporally aligned, instruction-following generation in both joint and unimodal settings while generalizing robustly to out-of-distribution scenarios. Across tasks, it substantially outperforms prior methods by a large margin and achieves performance comparable to Veo 3, offering a unified, scalable path toward next-generation audio-video synthesis.

音视频生成多模态扩散模型数据构建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。