arXiv:2507.15852cs.CVcs.AI2025-07被引 23

用概念构建提升视频对象分割,显著优于现有方法。

Advancing Complex Video Object Segmentation via Progressive Concept Construction

  • 基于大视觉语言模型构建物体中心的高层概念表示
  • 在新场景出现时触发概念注入,兼顾性能与效率
  • 在复杂语义场景下表现突出,适合高阶视觉理解任务

我们提出Segment Concept(SeC),一种以概念驱动的视频对象分割框架,将传统特征匹配转向逐步构建和利用高层、以物体为中心的表征。SeC利用大视觉语言模型(LVLMs)整合跨帧视觉线索,构建稳健的概念先验。为平衡语义推理与计算开销,仅在新场景出现时调用LVLM,于此时点注入概念级特征。为严格评估需高层次概念推理与鲁棒语义理解的VOS方法,我们引入语义复杂场景视频对象分割基准(SeCVOS)。SeCVOS包含160个手动标注的多场景视频,旨在挑战模型在显著外观变化与动态场景转换下的能力。实验表明,SeC在SeCVOS与标准VOS基准上均显著优于现有先进方法,包括SAM 2及其改进版本。尤其在SeCVOS上相较SAM 2.1提升11.8点,确立了概念感知视频分割的新标杆。

原文摘要 · Abstract (English)

We propose Segment Concept (SeC), a concept-driven video object segmentation (VOS) framework that shifts from conventional feature matching to the progressive construction and utilization of high-level, object-centric representations. SeC employs Large Vision-Language Models (LVLMs) to integrate visual cues across diverse frames, constructing robust conceptual priors. To balance semantic reasoning with computational overhead, SeC forwards the LVLMs only when a new scene appears, injecting concept-level features at those points. To rigorously assess VOS methods in scenarios demanding high-level conceptual reasoning and robust semantic understanding, we introduce the Semantic Complex Scenarios Video Object Segmentation benchmark (SeCVOS). SeCVOS comprises 160 manually annotated multi-scenario videos designed to challenge models with substantial appearance variations and dynamic scene transformations. Empirical evaluations demonstrate that SeC substantially outperforms state-of-the-art approaches, including SAM 2 and its advanced variants, on both SeCVOS and standard VOS benchmarks. In particular, SeC achieves an 11.8-point improvement over SAM 2.1 on SeCVOS, establishing a new state-of-the-art in concept-aware VOS.

视频分割概念构建大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。