无需训练即可提升视频生成音频的自然度与语义一致性。
Training-Free Multimodal Guidance for Video to Audio Generation
- 利用多模态嵌入体积实现视频、音频、文本统一对齐
- 在VGGSound和AudioCaps上显著提升音质与跨模态一致性
- 可无缝接入任意预训练音频扩散模型,轻量易用
视频到音频(V2A)生成旨在从无声视频中合成真实且语义一致的音频,适用于视频编辑、拟音设计及辅助多媒体等领域。现有方法要么需要在大规模配对数据集上进行昂贵的联合训练,要么依赖成对相似性,难以捕捉全局多模态一致性。本文提出一种全新的无训练多模态引导机制(MDG),基于视频、音频与文本嵌入所覆盖的体积,实现三者间的统一对齐。该方法作为轻量级即插即用信号,可直接应用于任何预训练音频扩散模型而无需重新训练。在VGGSound和AudioCaps数据集上的实验表明,相比基线方法,本方法在感知质量与多模态对齐方面均有显著提升,验证了联合多模态引导在V2A任务中的有效性。
原文摘要 · Abstract (English)
Video-to-audio (V2A) generation aims to synthesize realistic and semantically aligned audio from silent videos, with potential applications in video editing, Foley sound design, and assistive multimedia. Although the excellent results, existing approaches either require costly joint training on large-scale paired datasets or rely on pairwise similarities that may fail to capture global multimodal coherence. In this work, we propose a novel training-free multimodal guidance mechanism for V2A diffusion that leverages the volume spanned by the modality embeddings to enforce unified alignment across video, audio, and text. The proposed multimodal diffusion guidance (MDG) provides a lightweight, plug-and-play control signal that can be applied on top of any pretrained audio diffusion model without retraining. Experiments on VGGSound and AudioCaps demonstrate that our MDG consistently improves perceptual quality and multimodal alignment compared to baselines, proving the effectiveness of a joint multimodal guidance for V2A.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。