无需时空配对数据,用视觉提示定位特定声音源。
AV-SSAN: Audio-Visual Selective DoA Estimation through Explicit Multi-Band Semantic-Spatial Alignment
- 分频段对齐视听特征,实现跨实例目标定位
- 在新数据集上达16.59的平均误差和71.29%准确率
- 适合需要灵活声音定位的应用场景
视听声音源定位(AV-SSL)通过融合听觉与视觉线索估计声源位置。现有方法通常需要时空配对的音视频数据,且无法选择性定位特定目标。为此,我们提出跨实例视听定位(CI-AVL)新任务,即利用同一语义类别中不同实例的视觉提示来定位目标声音源,实现无需时空配对数据的定向定位。为解决该任务,我们提出基于多频带语义-空间对齐网络(MB-SSA Net)的AV-SSAN框架,将音频谱图分解为多个频段,分别与视觉提示对齐并优化空间线索以估计方向(DoA)。为支持本研究,我们构建了大规模数据集VGGSound-SSL,包含13,981段空间音频,涵盖296个类别,每段均配有视觉提示。AV-SSAN在该任务上取得16.59的平均绝对误差和71.29%的准确率,显著优于现有方法。代码与数据将公开。
原文摘要 · Abstract (English)
Audio-visual sound source localization (AV-SSL) estimates the position of sound sources by fusing auditory and visual cues. Current AV-SSL methodologies typically require spatially-paired audio-visual data and cannot selectively localize specific target sources. To address these limitations, we introduce Cross-Instance Audio-Visual Localization (CI-AVL), a novel task that localizes target sound sources using visual prompts from different instances of the same semantic class. CI-AVL enables selective localization without spatially paired data. To solve this task, we propose AV-SSAN, a semantic-spatial alignment framework centered on a Multi-Band Semantic-Spatial Alignment Network (MB-SSA Net). MB-SSA Net decomposes the audio spectrogram into multiple frequency bands, aligns each band with semantic visual prompts, and refines spatial cues to estimate the direction-of-arrival (DoA). To facilitate this research, we construct VGGSound-SSL, a large-scale dataset comprising 13,981 spatial audio clips across 296 categories, each paired with visual prompts. AV-SSAN achieves a mean absolute error of 16.59 and an accuracy of 71.29%, significantly outperforming existing AV-SSL methods. Code and data will be public.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。