用语音转文字+视觉语言模型,实现音频定位视频目标
3rd Place of MeViS-Audio Track of the 5th PVUW: VIRST-Audio
- 语音转文字后,用文本监督做视频目标分割
- 引入存在感知门控,减少错误分割,提升稳定性
- 在音频定位任务中表现可靠,适合实际应用
基于音频的指代视频对象分割(ARVOS)需要将音频查询与时间维度上的像素级对象掩码对齐,面临声学信号与时空视觉表征之间的桥接挑战。本文提出VIRST-Audio,一种基于预训练视频对象分割模型并融合视觉-语言架构的实用框架。不依赖音频专用训练,而是通过语音识别模块将输入音频转为文本,利用文本监督完成分割,实现从文本推理到音频驱动场景的有效迁移。为进一步提升鲁棒性,引入存在感知门控机制,估计目标对象是否存在于视频中,当目标不存在时抑制预测,减少幻觉掩码并稳定分割行为。我们在第五届PVUW挑战赛的MeViS-Audio赛道上评估该方法,VIRST-Audio取得第三名,证明其在音频指代视频分割任务中具备强泛化能力与可靠性能。
原文摘要 · Abstract (English)
Audio-based Referring Video Object Segmentation (ARVOS) requires grounding audio queries into pixel-level object masks over time, posing challenges in bridging acoustic signals with spatio-temporal visual representations. In this report, we present VIRST-Audio, a practical framework built upon a pretrained RVOS model integrated with a vision-language architecture. Instead of relying on audio-specific training, we convert input audio into text using an ASR module and perform segmentation using text-based supervision, enabling effective transfer from text-based reasoning to audio-driven scenarios. To improve robustness, we further incorporate an existence-aware gating mechanism that estimates whether the referred target object is present in the video and suppresses predictions when it is absent, reducing hallucinated masks and stabilizing segmentation behavior. We evaluate our approach on the MeViS-Audio track of the 5th PVUW Challenge, where VIRST-Audio achieves 3rd place, demonstrating strong generalization and reliable performance in audio-based referring video segmentation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。