用多模态联合训练,让视频生成高质量同步音频
MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis

- 联合训练视频与文本-音频数据,提升生成质量
- 8秒音频生成仅需1.23秒,参数量仅157M
- 兼顾视频/文本到音频生成,适合多场景应用
我们提出MMAudio框架,通过新型多模态联合训练,实现给定视频和可选文本条件下的高质量、高同步性音频合成。不同于仅依赖有限视频数据的单模态训练,MMAudio联合使用大规模易获取的文本-音频数据,学习生成语义对齐的高质量音频样本。此外,引入条件同步模块,在帧级对齐视频条件与音频潜在表示,增强音画同步。基于流匹配目标训练,MMAudio在公开模型中达到视频到音频生成的新基准,涵盖音频质量、语义对齐与音画同步三项指标,推理时间仅1.23秒生成8秒音频,参数量为157M。同时在文本到音频生成任务上表现优异,表明联合训练未损害单模态性能。代码与演示见:https://hkchengrex.github.io/MMAudio
原文摘要 · Abstract (English)
We propose to synthesize high-quality and synchronized audio, given video and optional text conditions, using a novel multimodal joint training framework MMAudio. In contrast to single-modality training conditioned on (limited) video data only, MMAudio is jointly trained with larger-scale, readily available text-audio data to learn to generate semantically aligned high-quality audio samples. Additionally, we improve audio-visual synchrony with a conditional synchronization module that aligns video conditions with audio latents at the frame level. Trained with a flow matching objective, MMAudio achieves new video-to-audio state-of-the-art among public models in terms of audio quality, semantic alignment, and audio-visual synchronization, while having a low inference time (1.23s to generate an 8s clip) and just 157M parameters. MMAudio also achieves surprisingly competitive performance in text-to-audio generation, showing that joint training does not hinder single-modality performance. Code and demo are available at: https://hkchengrex.github.io/MMAudio
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。