arXiv:2412.06209cs.CVcs.MM2024-12被引 7

用声音生成多样视觉画面,突破音画信息鸿沟

Sound2Vision: Generating Diverse Visuals from Audio through Cross-Modal Latent Alignment

论文配图:Sound2Vision: Generating Diverse Visuals from Audio through Cross-Modal Latent Alignment
图 1 · 摘自论文原文
  • 通过视觉信息增强音频特征并映射到图像潜在空间
  • 在VEGAS和VGGSound上显著优于现有方法
  • 支持简单波形或潜空间操作实现生成控制

声音如何描述周围世界?本文提出一种从多样的真实场景声音生成视觉画面的方法。该跨模态生成任务因听觉与视觉信号间存在巨大信息鸿沟而极具挑战。我们设计了一种模型,通过用视觉信息丰富音频特征,并将其翻译至视觉潜在空间,再输入预训练图像生成器以产出图像。为提升图像质量,利用声音源定位筛选具有强跨模态相关性的音视频对。实验表明,本方法在VEGAS和VGGSound数据集上的表现显著优于先前工作;且通过简单修改输入波形或潜空间即可实现生成过程的可控性。进一步分析显示,所学嵌入空间具备良好几何特性,证明该学习方式有效对齐音画信号。基于此,我们展示了方法对具体设计选择的无关性,验证了其通用性——可兼容多种模型架构及不同类型的音视频数据。

原文摘要 · Abstract (English)

How does audio describe the world around us? In this work, we propose a method for generating images of visual scenes from diverse in-the-wild sounds. This cross-modal generation task is challenging due to the significant information gap between auditory and visual signals. We address this challenge by designing a model that aligns audio-visual modalities by enriching audio features with visual information and translating them into the visual latent space. These features are then fed into the pre-trained image generator to produce images. To enhance image quality, we use sound source localization to select audio-visual pairs with strong cross-modal correlations. Our method achieves substantially better results on the VEGAS and VGGSound datasets compared to previous work and demonstrates control over the generation process through simple manipulations to the input waveform or latent space. Furthermore, we analyze the geometric properties of the learned embedding space and demonstrate that our learning approach effectively aligns audio-visual signals for cross-modal generation. Based on this analysis, we show that our method is agnostic to specific design choices, showing its generalizability by integrating various model architectures and different types of audio-visual data.

跨模态生成音画对齐图像生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。