发现音视频定位数据集存在视觉偏见,影响模型真实能力评估
Unveiling Visual Biases in Audio-Visual Localization Benchmarks
- 通过分析发现声音源常仅凭视觉就能识别
- 两个主流数据集上纯视觉模型反而优于音视频联合模型
- 提醒研究者需改进基准测试以真实评估跨模态学习能力
音视频源定位(AVSL)旨在定位视频中的声音来源。本文发现现有基准存在显著问题:声音源通常仅凭视觉线索即可轻易识别,这种现象称为视觉偏见。该偏见导致基准无法有效评估AVSL模型的真实性能。为验证这一假设,我们分析了两个代表性基准VGG-SS和EpicSounding-Object,结果显示纯视觉模型在这些数据集上的表现优于所有音视频联合基线模型。这表明当前的AVSL基准需要进一步优化,才能真正促进音视频协同学习。
原文摘要 · Abstract (English)
Audio-Visual Source Localization (AVSL) aims to localize the source of sound within a video. In this paper, we identify a significant issue in existing benchmarks: the sounding objects are often easily recognized based solely on visual cues, which we refer to as visual bias. Such biases hinder these benchmarks from effectively evaluating AVSL models. To further validate our hypothesis regarding visual biases, we examine two representative AVSL benchmarks, VGG-SS and EpicSounding-Object, where the vision-only models outperform all audiovisual baselines. Our findings suggest that existing AVSL benchmarks need further refinement to facilitate audio-visual learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。