提出自适应模态抑制方法,让模型根据语义自动忽略干扰信息
PRIMED: Adaptive Modality Suppression for Referring Audio-Visual Segmentation via Biased Competition

- 基于认知科学的偏见竞争理论,动态判断语义依赖的主要模态
- 在基准测试上达到当前最佳性能,显著提升定位准确率
- 适合需要精准跨模态理解的视频分析任务
指代式音视频分割(Ref-AVS)旨在根据视觉、听觉和文本提示,在视频帧中定位并分割目标对象。该任务具有挑战性,因为不同表达和场景中各模态的相关性各异,而现有方法通常将多模态线索视为同质输入进行融合或推理,易受无关或误导性模态影响。为此,我们提出PRIMED,受认知神经科学中偏见竞争理论启发,显式建模视觉感知与语言驱动先验调节,通过自适应模态抑制实现更精准的Ref-AVS。具体而言,模态先验解码器首先估计指代表达主要依赖音频、视觉或二者交互,生成模态先验以自适应引导高层注意力;令牌蒸馏模块从高层特征中提取紧凑的全局视觉令牌,并在具备竞争感知能力的跨模态融合模块间共享,提供分层全局上下文;此外,引入空间感知语义对齐损失,通过对比学习增强前景-背景区分能力。在Ref-AVS基准上的大量实验表明,PRIMED实现了当前最优的整体性能。
原文摘要 · Abstract (English)
Referring Audio-Visual Segmentation (Ref-AVS) seeks to localize and segment target objects in video frames based on visual, auditory, and textual referring cues. The task is challenging because the relevance of different modalities varies across referring expressions and scenes, while existing methods typically treat multimodal cues as homogeneous inputs for fusion, prompting, or reasoning, making them vulnerable to irrelevant or misleading modalities. To address this problem, we propose PRIMED, inspired by the biased competition theory in cognitive neuroscience, which explicitly models both visual perception and language-driven prior modulation, and enables more accurate Ref-AVS by adaptive modality suppression. Specifically, a Modality Prior Decoder first estimates whether the referring expression relies primarily on audio, vision, or their joint interaction, generating a modality prior to adaptively guide high-level attention. A Token Distiller further extracts compact global visual tokens from high-level features and shares them across Competition-aware Cross-modal Fusion modules to provide hierarchical global context. Additionally, we introduce a Spatial-Aware Semantic Alignment loss to further enhance foreground-background discrimination through contrastive learning. Extensive experiments on the Ref-AVS benchmark demonstrate that PRIMED achieves state-of-the-art overall performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。