arXiv:2507.20740cs.CV2025-07ICCV被引 6

通过隐式反事实学习,解决音视频分割中的模态偏差问题。

Implicit Counterfactual Learning for Audio-Visual Segmentation

  • 用多粒度隐式文本构建跨模态共享空间,缩小模态差异
  • 在三个公开数据集上达到最新最好性能,提升分割精度
  • 适合关注多模态融合与公平表示学习的研究者

音视频分割(AVS)旨在基于音频线索分割视频中的物体。现有方法主要提升交互效率,却忽视了模态表征差异与不平衡问题。为此,我们提出隐式反事实框架(ICF),实现无偏的跨模态理解。由于语义缺失,异构表征可能导致错误匹配,尤其在视觉内容模糊或多音源干扰的复杂场景中。我们引入多粒度隐式文本(MIT),涵盖视频、片段和帧级信息,作为桥梁建立模态共享空间,减少模态差距并提供先验引导。视觉内容信息量大,通常主导决策,导致音频特征被边缘化。为缓解知识偏好,提出语义反事实(SC),在潜在空间学习正交表示,生成多样反事实样本,避免因复杂功能设计或显式修改文本结构/属性带来的偏差。进一步提出协作分布感知对比学习(CDCL),融合事实-反事实与跨模态对比,对齐表征,促进一致性与解耦。在三个公开数据集上的大量实验验证,所提方法取得当前最优性能。

原文摘要 · Abstract (English)

Audio-visual segmentation (AVS) aims to segment objects in videos based on audio cues. Existing AVS methods are primarily designed to enhance interaction efficiency but pay limited attention to modality representation discrepancies and imbalances. To overcome this, we propose the implicit counterfactual framework (ICF) to achieve unbiased cross-modal understanding. Due to the lack of semantics, heterogeneous representations may lead to erroneous matches, especially in complex scenes with ambiguous visual content or interference from multiple audio sources. We introduce the multi-granularity implicit text (MIT) involving video-, segment- and frame-level as the bridge to establish the modality-shared space, reducing modality gaps and providing prior guidance. Visual content carries more information and typically dominates, thereby marginalizing audio features in the decision-making. To mitigate knowledge preference, we propose the semantic counterfactual (SC) to learn orthogonal representations in the latent space, generating diverse counterfactual samples, thus avoiding biases introduced by complex functional designs and explicit modifications of text structures or attributes. We further formulate the collaborative distribution-aware contrastive learning (CDCL), incorporating factual-counterfactual and inter-modality contrasts to align representations, promoting cohesion and decoupling. Extensive experiments on three public datasets validate that the proposed method achieves state-of-the-art performance.

音视频分割多模态学习反事实学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。