arXiv:2510.11063cs.CV2025-10被引 6

2025年大规模视频对象分割挑战赛报告,聚焦复杂场景下的鲁棒分割新进展。

LSVOS 2025 Challenge Report: Recent Advances in Complex Video Object Segmentation

  • 新增复杂视频分割赛道,模拟密集小目标与严重遮挡等真实场景
  • 引入新指标$J\&\dot{F}$,更有效评估跨尺度与消失对象的分割性能
  • 突出大模型与记忆感知传播技术在实际场景中的关键作用

本文概述了与ICCV 2025同期举办的第七届大规模视频对象分割(LSVOS)挑战赛。除传统的经典视频对象分割(VOS)和指代表达视频对象分割(RVOS)外,2025年新增复杂视频对象分割(MOSEv2)赛道。该赛道在以往基础上显著提升难度,引入更复杂的现实场景,如密集小物体、频繁消失/重现、严重遮挡、恶劣天气与光照等,对长期一致性与泛化能力提出更高要求。挑战赛沿用标准的${J}$、$F$和${J\&F}$指标评估VOS与RVOS,而MOSEv2采用${J\&\dot{F}}$作为主要排名指标,以更好衡量跨尺度及消失情况下的分割表现。报告总结了数据集与评测协议,分析了领先方案,并提炼出新兴趋势,如大语言模型(LLM)/多模态大模型(MLLM)组件和记忆感知传播机制的日益重要,旨在为未来真实场景下鲁棒、语言感知的视频分割指明方向。

原文摘要 · Abstract (English)

This report presents an overview of the 7th Large-scale Video Object Segmentation (LSVOS) Challenge held in conjunction with ICCV 2025. Besides the two traditional tracks of LSVOS that jointly target robustness in realistic video scenarios: Classic VOS (VOS), and Referring VOS (RVOS), the 2025 edition features a newly introduced track, Complex VOS (MOSEv2). Building upon prior insights, MOSEv2 substantially increases difficulty, introducing more challenging but realistic scenarios including denser small objects, frequent disappear/reappear events, severe occlusions, adverse weather and lighting, etc., pushing long-term consistency and generalization beyond curated benchmarks. The challenge retains standard ${J}$, $F$, and ${J\&F}$ metrics for VOS and RVOS, while MOSEv2 adopts ${J\&\dot{F}}$ as the primary ranking metric to better evaluate objects across scales and disappearance cases. We summarize datasets and protocols, highlight top-performing solutions, and distill emerging trends, such as the growing role of LLM/MLLM components and memory-aware propagation, aiming to chart future directions for resilient, language-aware video segmentation in the wild.

视频分割大模型挑战赛复杂场景

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。