统一视频文本生成音频,解决数据与模型竞争难题。
Omni2Sound: Towards Unified Video-Text-to-Audio Generation
- 构建47万对高质量音视频文本数据集SoundAtlas,实现精准对齐。
- 提出三阶段训练策略,单模型达成三项任务最佳性能。
- 适合多模态音频生成、跨模态理解研究者使用。
构建统一模型联合完成视频到音频(V2A)、文本到音频(T2A)及视频文本到音频(VT2A)生成,面临两大未解挑战:一是高质量音视频文本对齐数据稀缺,导致多模态条件间语义冲突;二是跨任务与任务内竞争,表现为V2A与T2A性能权衡及VT2A中的模态偏倚。为此,我们提出SoundAtlas,一个包含47万对样本的大规模数据集,通过视觉-语言压缩降低大模型视觉偏见,采用初级-高级代理交接机制降低5倍成本,并结合后处理过滤保障质量,实现语义丰富且时序精确的对齐标注。进一步,提出Omni2Sound,一种支持灵活输入的统一扩散模型,设计三阶段多任务渐进式训练流程,将跨任务竞争转为联合优化,缓解VT2A中模态偏倚,保持音视频对齐与离屏音频生成忠实度。最后构建VGGSound-Omni基准,包含挑战性离屏音频测试项。基于标准DiT骨干网络,Omni2Sound在单一模型内实现三项任务的统一最优表现,在异构输入条件下展现出强泛化能力。
原文摘要 · Abstract (English)
Training a unified model integrating video-to-audio (V2A), text-to-audio (T2A), and joint video-text-to-audio (VT2A) generation offers significant application flexibility, yet faces two unexplored foundational challenges: (1) the scarcity of high-quality audio captions with tight V-A-T alignment, leading to severe semantic conflict between multimodal conditions, and (2) cross-task and intra-task competition, manifesting as an adverse V2A-T2A performance trade-off and modality bias in the VT2A task. First, to address data scarcity, we introduce SoundAtlas, a large-scale dataset (470k pairs) that significantly outperforms existing benchmarks and even human experts in quality. Powered by a novel agentic pipeline, it integrates Vision-to-Language Compression to mitigate visual bias of MLLMs, a Junior-Senior Agent Handoff for a 5$\times$ cost reduction, and rigorous Post-hoc Filtering to ensure fidelity. Consequently, SoundAtlas delivers semantically rich and temporally detailed captions with tight V-A-T alignment. Second, we propose Omni2Sound, a unified VT2A diffusion model supporting flexible input modalities. To resolve the inherent cross-task and intra-task competition, we design a three-stage multi-task progressive training schedule that converts cross-task competition into joint optimization and mitigates modality bias in the VT2A task, maintaining both audio-visual alignment and off-screen audio generation faithfulness. Finally, we construct VGGSound-Omni, a comprehensive benchmark for unified evaluation, including challenging off-screen tracks. With a standard DiT backbone, Omni2Sound achieves unified SOTA performance across all three tasks within a single model, demonstrating strong generalization across benchmarks with heterogeneous input conditions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。