arXiv:2504.05178cs.CV2025-04被引 2

用多模型融合提升大模型在动作描述视频分割中的表现。

The 1st Solution for 4th PVUW MeViS Challenge: Unleashing the Potential of Large Multimodal Models for Referring Video Segmentation

  • 通过均匀采样帧增强模型对全视频理解。
  • 多专家模型集成降低单模型误判风险。
  • 在MeViS测试集上达61.98% J&F,获CVPR 2025挑战赛第一。

运动表达视频分割旨在根据输入的动作描述分割视频中的目标物体。与传统指代视频对象分割(RVOS)不同,它更关注动作信息及多对象表达,难度更高。近期,大模态模型(LMMs)凭借其强大的视觉-语言感知能力在RVOS中崭露头角。本文提出一种简单有效的推理优化方法,充分释放LMM在指代视频分割中的潜力。首先,以Sa2VA作为统一的基线模型,实现图像与视频的密集语义理解。其次,在推理过程中对视频帧进行均匀采样,提升模型对整个视频的理解能力。最后,通过融合多个专家模型的输出结果,缓解单一模型的错误预测。该方案在MeViS测试集上取得61.98% J&F指标,位列2025年CVPR PVUW挑战赛MeViS赛道第一名。

原文摘要 · Abstract (English)

Motion expression video segmentation is designed to segment objects in accordance with the input motion expressions. In contrast to the conventional Referring Video Object Segmentation (RVOS), it places emphasis on motion as well as multi-object expressions, making it more arduous. Recently, Large Multimodal Models (LMMs) have begun to shine in RVOS due to their powerful vision-language perception capabilities. In this work, we propose a simple and effective inference optimization method to fully unleash the potential of LMMs in referring video segmentation. Firstly, we use Sa2VA as our baseline, which is a unified LMM for dense grounded understanding of both images and videos. Secondly, we uniformly sample the video frames during the inference process to enhance the model's understanding of the entire video. Finally, we integrate the results of multiple expert models to mitigate the erroneous predictions of a single model. Our solution achieved 61.98% J&F on the MeViS test set and ranked 1st place in the 4th PVUW Challenge MeViS Track at CVPR 2025.

视频分割大模型多模态动作理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。