让模型学会在静音和噪声中定位声源,提升真实场景下的感知能力。
Learning from Silence and Noise for Visual Sound Source Localization
- 引入静音与噪声进行自监督训练,增强模型鲁棒性
- 新指标衡量视听特征在正负样本间的对齐与分离度
- 构建含负音频的IS3+数据集,适合真实复杂场景研究
视觉声源定位旨在根据视频音频检测发声源位置。尽管近期取得进展,现有方法仍存在两方面不足:一是面对静音、噪声及非可见声源等低音视频语义对应情况表现不佳,即存在负音频时性能下降;二是以往评估局限于单一可见声源的正样本场景,数据集与评价指标均未涵盖负音频。为此,本文提出三项贡献:首先,设计一种融合静音与噪声的训练策略,使自监督模型SSL-SaN在正样本上达领先性能,同时对负音频更具鲁棒性;其次,提出新度量,量化正负音视频对间视听特征的对齐与可分性权衡;第三,推出IS3+,一个扩展改进的含负音频合成数据集。相关数据、代码与评测工具已公开于https://xavijuanola.github.io/SSL-SaN/。
原文摘要 · Abstract (English)
Visual sound source localization is a fundamental perception task that aims to detect the location of sounding sources in a video given its audio. Despite recent progress, we identify two shortcomings in current methods: 1) most approaches perform poorly in cases with low audio-visual semantic correspondence such as silence, noise, and offscreen sounds, i.e. in the presence of negative audio; and 2) most prior evaluations are limited to positive cases, where both datasets and metrics convey scenarios with a single visible sound source in the scene. To address this, we introduce three key contributions. First, we propose a new training strategy that incorporates silence and noise, which improves performance in positive cases, while being more robust against negative sounds. Our resulting self-supervised model, SSL-SaN, achieves state-of-the-art performance compared to other self-supervised models, both in sound localization and cross-modal retrieval. Second, we propose a new metric that quantifies the trade-off between alignment and separability of auditory and visual features across positive and negative audio-visual pairs. Third, we present IS3+, an extended and improved version of the IS3 synthetic dataset with negative audio. Our data, metrics and code are available on the https://xavijuanola.github.io/SSL-SaN/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。