arXiv:2603.22732cs.CV2026-03中稿 · CVPR

用可学习的提示词提升音视频定位与分割效果

SOUPLE: Enhancing Audio-Visual Localization and Segmentation with Learnable Prompt Contexts

  • 引入可学习上下文令牌,融合视觉特征生成条件化提示
  • 在VGGSound等数据集上显著提升音视频定位与分割准确率
  • 适合研究多模态对齐与音频视觉理解的学者参考

大规模预训练图文模型具备强大的多模态表征能力,但将对比语言-图像预训练(CLIP)模型应用于音视频定位仍具挑战。用音频嵌入令牌([V_A])替代分类令牌([CLS])难以捕捉语义线索,且提示“一张[ V_A ]的照片”无法建立音频嵌入与上下文令牌间的有效关联。为此,我们提出声觉感知提示学习(SOUPLE),将固定提示替换为可学习的上下文令牌。这些令牌融合视觉特征,为掩码解码器生成条件化上下文,有效建立音频与视觉输入之间的语义对应关系。在VGGSound、SoundNet和AVSBench数据集上的实验表明,SOUPLE显著提升了音视频定位与分割性能。

原文摘要 · Abstract (English)

Large-scale pre-trained image-text models exhibit robust multimodal representations, yet applying the Contrastive Language-Image Pre-training (CLIP) model to audio-visual localization remains challenging. Replacing the classification token ([CLS]) with an audio-embedded token ([V_A]) struggles to capture semantic cues, and the prompt "a photo of a [V_A]" fails to establish meaningful connections between audio embeddings and context tokens. To address these issues, we propose Sound-aware Prompt Learning (SOUPLE), which replaces fixed prompts with learnable context tokens. These tokens incorporate visual features to generate conditional context for a mask decoder, effectively bridging semantic correspondence between audio and visual inputs. Experiments on VGGSound, SoundNet, and AVSBench demonstrate that SOUPLE improves localization and segmentation performance.

音视频定位多模态学习提示学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。