提出负音频测试集,揭示主流视觉声音定位模型存在关键缺陷。
A Critical Assessment of Visual Sound Source Localization Models Including Negative Audio
- 引入静音、噪声和屏幕外三种负音频场景评估模型
- 多数SOTA模型在负音频下无法正确调整预测结果
- 指出现有模型缺乏通用阈值,不适用于真实复杂场景
视觉声音源定位(VSSL)旨在通过融合音视频信息识别声音来源位置。尽管当前先进模型取得进展,但仍存在三类关键问题:一、评估仅针对图像中可见发声物体;二、常依赖发声物体尺寸的先验知识;三、缺乏真实场景下的通用定位阈值,因以往方法未考虑正负样本。本文构建新测试集与指标,涵盖无声、噪声和屏幕外三种负音频情况。分析显示,多个SOTA模型未能根据音频输入合理调整输出,表明其可能未真正利用音频信息。同时,我们分析了正负音频下音视频相似度图的最大值范围,发现多数模型判别能力不足,无法设定无需先验信息(如物体大小或可见性)的通用阈值。
原文摘要 · Abstract (English)
The task of Visual Sound Source Localization (VSSL) involves identifying the location of sound sources in visual scenes, integrating audio-visual data for enhanced scene understanding. Despite advancements in state-of-the-art (SOTA) models, we observe three critical flaws: i) The evaluation of the models is mainly focused in sounds produced by objects that are visible in the image, ii) The evaluation often assumes a prior knowledge of the size of the sounding object, and iii) No universal threshold for localization in real-world scenarios is established, as previous approaches only consider positive examples without accounting for both positive and negative cases. In this paper, we introduce a novel test set and metrics designed to complete the current standard evaluation of VSSL models by testing them in scenarios where none of the objects in the image corresponds to the audio input, i.e. a negative audio. We consider three types of negative audio: silence, noise and offscreen. Our analysis reveals that numerous SOTA models fail to appropriately adjust their predictions based on audio input, suggesting that these models may not be leveraging audio information as intended. Additionally, we provide a comprehensive analysis of the range of maximum values in the estimated audio-visual similarity maps, in both positive and negative audio cases, and show that most of the models are not discriminative enough, making them unfit to choose a universal threshold appropriate to perform sound localization without any a priori information of the sounding object, that is, object size and visibility.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。