arXiv:2512.20117cs.CVcs.SD2025-12中稿 · ECCV被引 3

解决音视频分割中多声源混杂与对齐不准问题,提升细粒度定位能力

Delayed Bidirectional Alignment via Disentangled Audio Semantics for Audio-Visual Segmentation

  • 用可学习查询和音频原型记忆库解耦音频语义,减少多源混淆
  • 引入延迟双模态交叉注意力,增强音视频对齐鲁棒性
  • 在多个基准上达到最优性能,适合复杂真实场景的音视频分析

音视频分割(AVS)旨在通过融合听觉与视觉线索,在像素级定位发声物体。然而,现有方法常受多源混杂和音视频错位影响,导致对声音响亮或视觉显著物体产生偏好,忽视细微或共现声源。为此,本文提出DDAVS:基于解耦音频语义的延迟双向对齐音视频分割框架。为缓解多源混杂,DDAVS采用可学习查询提取音频语义,并将其锚定在由音频原型记忆库构建的结构化语义空间中,通过对比学习优化以增强判别力与鲁棒性。为缓解音视频错位,引入延迟模态交互的双重交叉注意力机制,提升多模态对齐稳定性。在AVS-Objects和VPO基准上的大量实验表明,DDAVS在单源、多源及多类多实例场景下均达到当前最优性能,验证了其在复杂真实场景下的有效性与泛化能力。

原文摘要 · Abstract (English)

Audio-Visual Segmentation (AVS) aims to localize sound-producing objects at the pixel level by integrating auditory and visual cues. However, existing methods often struggle with multi-source entanglement and audio-visual misalignment, leading to a dominance bias toward acoustically or visually salient objects (i.e., louder or larger ones) at the expense of subtler or co-occurring sources. To address these challenges, we propose DDAVS: Delayed Bidirectional Alignment via Disentangled Audio Semantics for Audio-Visual Segmentation. To mitigate multi-source entanglement, DDAVS employs learnable queries to extract audio semantics and anchor them within a structured semantic space derived from an audio prototype memory bank. This process is further optimized through contrastive learning to enhance discriminability and robustness. To alleviate audio-visual misalignment, DDAVS introduces dual cross attention with delayed modality interaction, improving the robustness of multimodal alignment. Extensive experiments on the AVS-Objects and VPO benchmarks demonstrate that DDAVS achieves state-of-the-art performance across single-source, multi-source, and multi-class multi-instance scenarios. These results validate the effectiveness and generalization ability of our framework under challenging real-world audio-visual segmentation conditions. Project page: https://trilarflagz.github.io/DDAVS-page/

音视频分割多模态对齐解耦表示跨模态学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。