arXiv:2501.05413cs.SDcs.CV2025-01

用视觉生成声音,突破音画配对数据瓶颈

Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation

  • 通过视觉语言模型检索,人工合成音画配对数据
  • 训练出性能媲美顶尖的音频转图像模型
  • 实现语义混合、音量校准等隐含声学建模能力

训练音频转图像生成模型需要大量语义对齐的音画配对数据。这类数据通常从真实视频中采集,依赖天然的跨模态对应关系。本文提出,强制要求真实音画对应不仅不必要,还会严重限制数据规模、质量与多样性,影响现代生成模型的性能。为此,我们设计了一种可扩展的图像声化框架:从多种高质量单模态来源获取实例,利用现代视觉-语言模型的推理能力,通过检索实现人工配对。实验表明,使用这些声化图像训练的音频转图像模型性能达到当前最优水平。通过一系列消融实验,我们发现模型隐式具备语义混合、插值、音量校准及混响相关的声学空间建模等能力,有效指导图像生成。

原文摘要 · Abstract (English)

Training audio-to-image generative models requires an abundance of diverse audio-visual pairs that are semantically aligned. Such data is almost always curated from in-the-wild videos, given the cross-modal semantic correspondence that is inherent to them. In this work, we hypothesize that insisting on the absolute need for ground truth audio-visual correspondence, is not only unnecessary, but also leads to severe restrictions in scale, quality, and diversity of the data, ultimately impairing its use in the modern generative models. That is, we propose a scalable image sonification framework where instances from a variety of high-quality yet disjoint uni-modal origins can be artificially paired through a retrieval process that is empowered by reasoning capabilities of modern vision-language models. To demonstrate the efficacy of this approach, we use our sonified images to train an audio-to-image generative model that performs competitively against state-of-the-art. Finally, through a series of ablation studies, we exhibit several intriguing auditory capabilities like semantic mixing and interpolation, loudness calibration and acoustic space modeling through reverberation that our model has implicitly developed to guide the image generation process.

音频生成多模态生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。