arXiv:2509.17537cs.CV2025-09被引 5

用大模型生成语义标记,精准定位视频中被语言描述的物体。

SimToken: A Simple Baseline for Referring Audio-Visual Segmentation

  • 用多模态大模型生成代表目标的紧凑语义标记。
  • 在基准数据集上超越现有方法,实现更优分割精度。
  • 适合做跨模态理解与细粒度定位研究者参考。

指称式音视频分割(Ref-AVS)旨在根据包含音频、视觉和文本信息的自然语言描述,在视频中分割出特定对象。该任务在跨模态推理和细粒度目标定位方面面临重大挑战。本文提出一种简单框架 SimToken,将多模态大语言模型(MLLM)与 Segment Anything Model(SAM)结合。MLLM 被引导生成一个特殊语义标记,代表被提及的对象。该紧凑标记融合了多模态上下文信息,作为提示引导 SAM 在视频帧间完成分割。为进一步提升语义学习,我们引入一种新颖的目标一致语义对齐损失,使指向同一对象的不同表达所生成的标记嵌入保持一致。在 Ref-AVS 基准上的实验表明,该方法优于现有方法。

原文摘要 · Abstract (English)

Referring Audio-Visual Segmentation (Ref-AVS) aims to segment specific objects in videos based on natural language expressions involving audio, vision, and text information. This task poses significant challenges in cross-modal reasoning and fine-grained object localization. In this paper, we propose a simple framework, SimToken, that integrates a multimodal large language model (MLLM) with the Segment Anything Model (SAM). The MLLM is guided to generate a special semantic token representing the referred object. This compact token, enriched with contextual information from all modalities, acts as a prompt to guide SAM to segment objectsacross video frames. To further improve semantic learning, we introduce a novel target-consistent semantic alignment loss that aligns token embeddings from different expressions but referring to the same object. Experiments on the Ref-AVS benchmark demonstrate that our approach achieves superior performance compared to existing methods.

音视频分割多模态大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。