arXiv:2506.18557cs.CV2025-06CVPR被引 10

用多模态大模型区分发声与静止物体,提升复杂场景音源定位准确率。

Object-aware Sound Source Localization via Audio-Visual Scene Understanding

  • 引入MLLM生成显式语义差异描述,区分发声与无声物体
  • 在MUSIC和VGGSound上显著优于现有方法,单/多音源定位均提升
  • 适合需要高精度音视频定位的智能监控、机器人感知等场景

音频-视觉音源定位任务旨在通过融合视觉与音频线索,在视觉场景中精确定位发声物体。然而,现有方法在复杂场景中难以准确识别发声物体,尤其当存在视觉相似的静止物体时。这一局限主要源于其依赖简单的音视频对应关系,未能捕捉发声与静止物体之间的细粒度语义差异。为此,我们提出一种新型音源定位框架,利用多模态大语言模型(MLLM)生成详细上下文信息,明确区分发声前景物体与静止背景物体。为有效融合这些信息,我们引入两种新损失函数:物体感知对比对齐(OCA)损失与物体区域隔离(ORI)损失。在MUSIC和VGGSound数据集上的大量实验表明,该方法在单源与多源定位场景中均显著优于现有方法。代码与生成的详细上下文信息已开源:https://github.com/VisualAIKHU/OA-SSL。

原文摘要 · Abstract (English)

Audio-visual sound source localization task aims to spatially localize sound-making objects within visual scenes by integrating visual and audio cues. However, existing methods struggle with accurately localizing sound-making objects in complex scenes, particularly when visually similar silent objects coexist. This limitation arises primarily from their reliance on simple audio-visual correspondence, which does not capture fine-grained semantic differences between sound-making and silent objects. To address these challenges, we propose a novel sound source localization framework leveraging Multimodal Large Language Models (MLLMs) to generate detailed contextual information that explicitly distinguishes between sound-making foreground objects and silent background objects. To effectively integrate this detailed information, we introduce two novel loss functions: Object-aware Contrastive Alignment (OCA) loss and Object Region Isolation (ORI) loss. Extensive experimental results on MUSIC and VGGSound datasets demonstrate the effectiveness of our approach, significantly outperforming existing methods in both single-source and multi-source localization scenarios. Code and generated detailed contextual information are available at: https://github.com/VisualAIKHU/OA-SSL.

音源定位多模态大模型视觉理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。