arXiv:2509.23838cs.CV2025-09

用概念理解提升视频目标分割鲁棒性,零样本通过复杂挑战赛

2nd Place Report of MOSEv2 Challenge 2025: Concept Guided Video Object Segmentation via SeC

  • 基于大视觉语言模型建立目标概念理解,替代传统外观匹配
  • 零样本在MOSEv2数据集上达到39.7的JF score,排名第二
  • 适合需要强泛化能力的视频分割场景,如遮挡与视角剧变

半监督视频目标分割旨在通过首帧掩码对视频序列中的指定目标进行分割。以往方法严重依赖外观模式匹配,在剧烈视觉变化、遮挡和场景切换下表现不佳,根源在于缺乏对目标的高层概念理解。最近提出的SeC框架利用大视觉语言模型(LVLM)建立对象深层语义理解,提升了分割持续性。本文评估其在挑战性coMplex video Object SEgmentation v2(MOSEv2)数据集上的零样本性能。未在训练集上微调,SeC在测试集上取得39.7的JF score,位列第7届大规模视频目标分割挑战赛复杂VOS赛道第2名。

原文摘要 · Abstract (English)

Semi-supervised Video Object Segmentation aims to segment a specified target throughout a video sequence, initialized by a first-frame mask. Previous methods rely heavily on appearance-based pattern matching and thus exhibit limited robustness against challenges such as drastic visual changes, occlusions, and scene shifts. This failure is often attributed to a lack of high-level conceptual understanding of the target. The recently proposed Segment Concept (SeC) framework mitigated this limitation by using a Large Vision-Language Model (LVLM) to establish a deep semantic understanding of the object for more persistent segmentation. In this work, we evaluate its zero-shot performance on the challenging coMplex video Object SEgmentation v2 (MOSEv2) dataset. Without any fine-tuning on the training set, SeC achieved 39.7 \JFn on the test set and ranked 2nd place in the Complex VOS track of the 7th Large-scale Video Object Segmentation Challenge.

视频分割概念理解零样本视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。