arXiv:2603.14203cs.CV2026-03中稿 · IEEE Transactions …被引 2

提升音视频分割在复杂场景下的抗噪能力与跨模态交互效果

Selective Noise Suppression and Discriminative Mutual Interaction for Robust Audio-Visual Segmentation

  • 设计选择性降噪模块,精准增强有效音频特征
  • 提出判别式融合策略,在多音源场景中实现更优表现
  • 适合研究音视频协同建模或实际应用中的鲁棒性问题

在动态视觉场景中准确捕捉并分割发声物体,对音视频分割(AVS)任务至关重要。尽管已有显著进展,但音视频模态间的交互仍需深入探索。本文针对两个核心问题:如何有效抑制音频噪声并增强相关音频信息?如何实现音视频模态间的判别性交互?为此提出SDAVS方法,包含选择性抗噪处理器(SNRP)模块与判别式音视频互融(DAMF)策略。SNRP通过选择性强化相关听觉线索,缓解噪声干扰;DAMF则确保音视频表示的一致性。实验表明,该方法在基准AVS数据集上取得领先性能,尤其在多音源和复杂场景中表现优异。代码与模型已开源。

原文摘要 · Abstract (English)

The ability to capture and segment sounding objects in dynamic visual scenes is crucial for the development of Audio-Visual Segmentation (AVS) tasks. While significant progress has been made in this area, the interaction between audio and visual modalities still requires further exploration. In this work, we aim to answer the following questions: How can a model effectively suppress audio noise while enhancing relevant audio information? How can we achieve discriminative interaction between the audio and visual modalities? To this end, we propose SDAVS, equipped with the Selective Noise-Resilient Processor (SNRP) module and the Discriminative Audio-Visual Mutual Fusion (DAMF) strategy. The proposed SNRP mitigates audio noise interference by selectively emphasizing relevant auditory cues, while DAMF ensures more consistent audio-visual representations. Experimental results demonstrate that our proposed method achieves state-of-the-art performance on benchmark AVS datasets, especially in multi-source and complex scenes. \textit{The code and model are available at https://github.com/happylife-pk/SDAVS}.

音视频分割跨模态融合降噪

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。