不训练提升视频目标分割效果,用语言检查和关键帧采样减少误判。
Enhancing Sa2VA for Referent Video Object Segmentation: 2nd Solution for 7th LSVOS RVOS Track
- 引入语言-视频校验器,确认描述内容是否真实出现
- 自适应选取关键帧,更好捕捉物体出现与长期时序信息
- 零训练实现64.14%准确率,获ICCV 2025第二名
参照性视频对象分割(RVOS)旨在分割视频中所有符合自然语言描述的对象,弥合视觉与语言理解之间的差距。近期工作如Sa2VA结合大语言模型(LLMs)与SAM~2,利用LLMs强大的视频推理能力引导分割。本文提出一种无需训练的框架,显著提升Sa2VA在RVOS任务上的表现。方法包含两个核心组件:(1) 视频-语言校验器,显式验证查询中的主体与动作是否实际出现在视频中,从而降低误报;(2) 关键帧采样器,自适应选择具有信息量的帧,以更好捕捉物体初始出现及长程时序上下文。无需额外训练,该方法在MeViS测试集上达到64.14%的J&F分数,在ICCV 2025第7届LSVOS挑战赛的RVOS赛道中位列第二。
原文摘要 · Abstract (English)
Referential Video Object Segmentation (RVOS) aims to segment all objects in a video that match a given natural language description, bridging the gap between vision and language understanding. Recent work, such as Sa2VA, combines Large Language Models (LLMs) with SAM~2, leveraging the strong video reasoning capability of LLMs to guide video segmentation. In this work, we present a training-free framework that substantially improves Sa2VA's performance on the RVOS task. Our method introduces two key components: (1) a Video-Language Checker that explicitly verifies whether the subject and action described in the query actually appear in the video, thereby reducing false positives; and (2) a Key-Frame Sampler that adaptively selects informative frames to better capture both early object appearances and long-range temporal context. Without any additional training, our approach achieves a J&F score of 64.14% on the MeViS test set, ranking 2nd place in the RVOS track of the 7th LSVOS Challenge at ICCV 2025.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。