arXiv:2411.15447cs.MMcs.CV2024-11AAAI被引 3

让声音生成更精准:通过识别声源位置和特性提升音效真实感

Gotta Hear Them All: Towards Sound Source Aware Audio Generation

  • 基于视觉检测与跨模态翻译,定位场景中每个声源并建模其特征
  • 在图像到音频任务中达到当前最优效果,局部音频相关性提升显著
  • 适合需要精细音效控制的多媒体生成、影视配乐等场景

音频合成在多媒体应用中前景广阔。近期方法可从描述音频场景的图像或文本生成相应音频,但沉浸感与表现力受限。主要问题在于现有方法仅依赖全局场景,忽略局部声源细节。为此,我们提出声源感知音频生成器SS2A:它通过视觉检测识别多模态声源,利用跨模态翻译提取特征;对比学习构建跨模态声源流形(CMSS)以语义区分各声源;最后将CMSS语义注意力融合为丰富音频表示,由预训练音频生成器输出。为建模该流形,我们构建了新数据集VGGS3(基于VGGSound),并设计声源匹配评分以量化局部音频相关性。实验表明,显式建模声源使SS2A在多种图像到音频任务中达到最先进性能。定性结果显示,可通过组合视觉、文本与音频条件实现直观合成控制。此外,结合简单的时间聚合机制,该方法在视频到音频任务中也表现出色。

原文摘要 · Abstract (English)

Audio synthesis has broad applications in multimedia. Recent advancements have made it possible to generate relevant audios from inputs describing an audio scene, such as images or texts. However, the immersiveness and expressiveness of the generation are limited. One possible problem is that existing methods solely rely on the global scene and overlook details of local sounding objects (i.e., sound sources). To address this issue, we propose a Sound Source-Aware Audio (SS2A) generator. SS2A is able to locally perceive multimodal sound sources from a scene with visual detection and cross-modality translation. It then contrastively learns a Cross-Modal Sound Source (CMSS) Manifold to semantically disambiguate each source. Finally, we attentively mix their CMSS semantics into a rich audio representation, from which a pretrained audio generator outputs the sound. To model the CMSS manifold, we curate a novel single-sound-source visual-audio dataset VGGS3 from VGGSound. We also design a Sound Source Matching Score to clearly measure localized audio relevance. With the effectiveness of explicit sound source modeling, SS2A achieves state-of-the-art performance in extensive image-to-audio tasks. We also qualitatively demonstrate SS2A's ability to achieve intuitive synthesis control by compositing vision, text, and audio conditions. Furthermore, we show that our sound source modeling can achieve competitive video-to-audio performance with a straightforward temporal aggregation mechanism.

音频生成声源建模跨模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。