arXiv:2502.00358cs.SDcs.AI2025-02AAAI被引 6

现有音视频分割模型依赖视觉显著性,难以真实识别发声物体。

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects?

  • 通过引入负样本训练和分类器引导相似性学习,增强音频感知。
  • 在无声、环境噪音等场景下,模型误报率接近零,性能显著提升。
  • 新基准测试揭示主流方法普遍存在视觉主导偏差,适合研究鲁棒性者参考。

与传统视觉分割不同,音视频分割(AVS)要求模型不仅识别并分割物体,还需判断其是否为声源。近期基于Transformer和SAM等基础模型的AVS方法在标准基准上表现优异,但关键问题仍存:这些模型是否真正融合音视频线索来分割发声物体?本文系统研究了鲁棒性AVS中的这一问题,发现当前方法存在根本性偏差——主要依据视觉显著性生成分割掩码,忽视音频上下文。这导致声音缺失或无关时预测不可靠。为此,我们提出AVSBench-Robust基准,涵盖静音、环境噪声及离屏声音等多种负音频场景。同时,设计一种结合平衡训练与负样本、分类器引导相似性学习的简单有效方法。大量实验表明,现有最先进方法在负音频条件下持续失效,凸显视觉偏见普遍性;而我们的方法在标准指标与鲁棒性度量上均实现显著提升,保持近乎完美的误报率,同时维持高质量分割性能。

原文摘要 · Abstract (English)

Unlike traditional visual segmentation, audio-visual segmentation (AVS) requires the model not only to identify and segment objects but also to determine whether they are sound sources. Recent AVS approaches, leveraging transformer architectures and powerful foundation models like SAM, have achieved impressive performance on standard benchmarks. Yet, an important question remains: Do these models genuinely integrate audio-visual cues to segment sounding objects? In this paper, we systematically investigate this issue in the context of robust AVS. Our study reveals a fundamental bias in current methods: they tend to generate segmentation masks based predominantly on visual salience, irrespective of the audio context. This bias results in unreliable predictions when sounds are absent or irrelevant. To address this challenge, we introduce AVSBench-Robust, a comprehensive benchmark incorporating diverse negative audio scenarios including silence, ambient noise, and off-screen sounds. We also propose a simple yet effective approach combining balanced training with negative samples and classifier-guided similarity learning. Our extensive experiments show that state-of-theart AVS methods consistently fail under negative audio conditions, demonstrating the prevalence of visual bias. In contrast, our approach achieves remarkable improvements in both standard metrics and robustness measures, maintaining near-perfect false positive rates while preserving highquality segmentation performance.

音视频分割鲁棒性视觉偏差

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。