让模型同时识别和定位画面中的语音与非语音声音
Seeing Speech and Sound: Distinguishing and Locating Audios in Visual Scenes
- 采用混音分离框架,联合学习音频与视觉的对应关系
- 在混合音频场景下,定位准确率显著优于现有方法
- 适合需要精准音频定位的多模态应用开发
我们提出一个统一模型,可同时在视觉场景中定位语音和非语音声音,解决了现有音频-视觉定位模型的局限性。当前方法通常只能分别处理语音或非语音声音,或虽能处理但仅按顺序进行,无法应对真实世界中常见的混合声音。我们的方法引入‘混音-分离’框架,结合视听对齐目标,利用混合音频联合学习对应关系与解耦表示。通过该机制,模型能生成区分度高的音频嵌入,实现对混合声源的有效解耦与定位。此外,我们构建了新数据集以评估混合音频的同步定位能力,结果表明本模型性能超越先前方法。该方法在标准分割与跨模态检索任务上也达到相当或更优表现,验证了混音分离策略的优势。
原文摘要 · Abstract (English)
We present a unified model capable of simultaneously grounding both spoken language and non-speech sounds within a visual scene, addressing key limitations in current audio-visual grounding models. Existing approaches are typically limited to handling either speech or non-speech sounds independently, or at best, together but sequentially without mixing. This limitation prevents them from capturing the complexity of real-world audio sources that are often mixed. Our approach introduces a 'mix-and-separate' framework with audio-visual alignment objectives that jointly learn correspondence and disentanglement using mixed audio. Through these objectives, our model learns to produce distinct embeddings for each audio type, enabling effective disentanglement and grounding across mixed audio sources. Additionally, we created a new dataset to evaluate simultaneous grounding of mixed audio sources, demonstrating that our model outperforms prior methods. Our approach also achieves comparable or better performance in standard segmentation and cross-modal retrieval tasks, highlighting the benefits of our mix-and-separate approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。