用运动流和文本提示协同提升音视频语义分割精度
How Do Optical Flow and Textual Prompts Collaborate to Assist in Audio-Visual Semantic Segmentation?
- 结合光流与文本提示,分步引导分割过程
- 在动态与静态声源场景下均实现更精准的像素级分割
- 适合需要跨模态理解的多媒体分析任务
音视频语义分割(AVSS)超越了传统音视频分割(AVS)仅识别发声物体位置的任务,要求对音视频场景进行语义理解。本文提出一种新框架 SSP,将光流与文本提示协同用于分割。针对移动声源,利用光流捕捉运动动态以提供时序上下文;针对静止声源(如闹钟),引入两类文本提示:一类标注发声物体类别,另一类描述场景整体。同时设计视觉-文本对齐模块(VTA)促进跨模态融合,提升语义一致性。训练中采用后掩码策略,强化模型对光流图的建模能力。实验表明,SSP优于现有方法,在多种场景下实现高效且精确的分割。
原文摘要 · Abstract (English)
Audio-visual semantic segmentation (AVSS) represents an extension of the audio-visual segmentation (AVS) task, necessitating a semantic understanding of audio-visual scenes beyond merely identifying sound-emitting objects at the visual pixel level. Contrary to a previous methodology, by decomposing the AVSS task into two discrete subtasks by initially providing a prompted segmentation mask to facilitate subsequent semantic analysis, our approach innovates on this foundational strategy. We introduce a novel collaborative framework, \textit{S}tepping \textit{S}tone \textit{P}lus (SSP), which integrates optical flow and textual prompts to assist the segmentation process. In scenarios where sound sources frequently coexist with moving objects, our pre-mask technique leverages optical flow to capture motion dynamics, providing essential temporal context for precise segmentation. To address the challenge posed by stationary sound-emitting objects, such as alarm clocks, SSP incorporates two specific textual prompts: one identifies the category of the sound-emitting object, and the other provides a broader description of the scene. Additionally, we implement a visual-textual alignment module (VTA) to facilitate cross-modal integration, delivering more coherent and contextually relevant semantic interpretations. Our training regimen involves a post-mask technique aimed at compelling the model to learn the diagram of the optical flow. Experimental results demonstrate that SSP outperforms existing AVS methods, delivering efficient and precise segmentation results.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。