arXiv:2603.27241cs.CV2026-03

通过目标存在性验证机制,提升视频对象分割对动态描述的响应能力。

SaSaSaSa2VA: 2nd Place of the 5th PVUW MeViS-Text Track

  • 引入目标存在性感知验证机制,增强模型对运动描述的理解。
  • 在MeViS-Text Track中取得89.19分,排名第二。
  • 方法简洁有效,特别适合处理含运动表达的视频指代任务。

指代视频对象分割(RVOS)通常基于静态文本线索定位目标。MeViS基准通过引入以运动为中心的表达(指代与推理类运动表达)及无目标查询,扩展了该任务。在已有输入帧增加和[SEG]标记强化Sa2VA骨干网络的基础上,本文提出一种简单而有效的目标存在性感知验证机制,形成改进版Still Awesome SaSaSa2VA(SaSaSaSa2VA)。尽管方法简洁,其在第五次PVUW挑战赛(MeViS-Text Track)中取得89.19分,位列第二。定量结果与消融实验表明,该存在性感知验证策略足以充分释放模型在运动中心指代任务中的性能潜力。

原文摘要 · Abstract (English)

Referring video object segmentation (RVOS) commonly grounds targets in videos based on static textual cues. MeViS benchmark extends this by incorporating motion-centric expressions (referring & reasoning motion expressions) and introducing no-target queries. Extending SaSaSa2VA, where increased input frames and [SEG] tokens already strengthen the Sa2VA backbone, we adopt a simple yet effective target existence-aware verification mechanism, leading to Still Awesome SaSaSa2VA (SaSaSaSa2VA). Despite its simplicity, the method achieves a final score of 89.19 in the 5th PVUW Challenge (MeViS-Text Track), securing 2nd place. Both quantitative results and ablations suggest that this existence-aware verification strategy is sufficient to unlock strong performance on motion-centric referring tasks.

视频分割指代理解运动感知多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。