arXiv:2506.01558cs.CV2025-06CVPR被引 24

用文本音频视觉融合提示,让SAM2在多模态场景中精准持续分割目标。

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes

  • 将文、音、视三模态融合为可学习的提示令牌,驱动SAM2进行像素级分割。
  • 在Ref-AVS基准上性能领先当前最优8.5%(J&F),实现帧间一致分割。
  • 适合研究多模态理解、视频分割与智能交互的开发者和研究者。

参考音频-视觉分割(Ref-AVS)旨在语言辅助的音视频场景(LAVS)中实现像素级场景理解,要求模型从视频中连续分割出由文本和音频所指的对象。以往双模态方法因缺少第三模态而失败,现有三模态方法则存在时空不一致问题,导致不同帧间目标漂移。本文提出新框架SAM2-LOVE,将文本、音频与视觉表示融合为可学习的提示令牌,用于引导并对齐SAM2,在LAVS中实现Ref-AVS。技术上,设计多模态融合模块提升理解能力,并引入令牌传播与累积策略,在不遗忘历史信息的前提下增强时空一致性。大量实验表明,SAM2-LOVE在Ref-AVS基准上以8.5%的显著优势超越当前最优(J&F指标),验证了各组件的简洁性与有效性。代码将公开。

原文摘要 · Abstract (English)

Reference Audio-Visual Segmentation (Ref-AVS) aims to provide a pixel-wise scene understanding in Language-aided Audio-Visual Scenes (LAVS). This task requires the model to continuously segment objects referred to by text and audio from a video. Previous dual-modality methods always fail due to the lack of a third modality and the existing triple-modality method struggles with spatio-temporal consistency, leading to the target shift of different frames. In this work, we introduce a novel framework, termed SAM2-LOVE, which integrates textual, audio, and visual representations into a learnable token to prompt and align SAM2 for achieving Ref-AVS in the LAVS. Technically, our approach includes a multimodal fusion module aimed at improving multimodal understanding of SAM2, as well as token propagation and accumulation strategies designed to enhance spatio-temporal consistency without forgetting historical information. We conducted extensive experiments to demonstrate that SAM2-LOVE outperforms the SOTA by 8.5\% in $\mathcal{J\&F}$ on the Ref-AVS benchmark and showcase the simplicity and effectiveness of the components. Our code will be available here.

多模态视频分割SAM2语音理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。