arXiv:2504.00476cs.CV2025-04被引 2

用简单推理优化提升多模态大模型在视频目标分割上的表现

4th PVUW MeViS 3rd Place Report: Sa2VA

  • 通过扩大关键帧范围改进测试时推理,无需额外训练
  • 在MeViS数据集上取得第三名,超越现有基准
  • 适合关注多模态大模型应用与视频理解的研究者

指称视频对象分割(RVOS)是一项挑战性任务,要求模型根据语言描述分割视频中的目标物体。MeViS是近期提出的数据集,包含目标物体的运动表达,相比现有基准更具挑战性。与此同时,指称表达任务的新趋势是采用多模态大语言模型(MLLM)以实现更优的图像与文本对齐。本报告表明,仅通过在更强的MLLM上对测试时推理方法进行简单修改,即可在MeViS上获得更优结果。具体而言,我们采用近期提出的统一模型Sa2VA,该模型可实现图像与视频的密集语义理解。通过扩大关键帧范围,在不进行任何额外训练的情况下,我们在第四届PVUW研讨会中取得第三名。

原文摘要 · Abstract (English)

Referring video object segmentation (RVOS) is a challenging task that requires the model to segment the object in a video given the language description. MeViS is a recently proposed dataset that contains motion expressions of the target objects, leading to a challenging benchmark, compared with existing RVOS benchmarks. On the other hand, for referring expression tasks, a new trend is to adopt multi-modal large language model (MLLM) to achieve better image and text alignment. In this report, we show that with a simple modification to the test time inference method on stronger MLLMs, we can lead to stronger results on MeVIS. In particular, we adopt the recent method Sa2VA, a unified model for dense grounded understanding of both images and videos. By enlarging the scope of key frames, without any further training, we can achieve the 3rd place in the 4th PVUW workshop.

视频分割多模态大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。