8th LSVOS挑战赛报告:多模态视频目标分割新范式
Report of the 8th LSVOS Challenge: Complex and Multimodal Video Object Segmentation

- 融合多模态提示与目标感知记忆,实现复杂场景下的精准分割
- 在三个赛道上最优模型性能超越基线30%以上
- 适合关注多模态视频理解与模块化分割系统的研究者
本报告总结了与ECCV 2026同期举办的第八届大规模视频对象分割(LSVOS)挑战赛。挑战赛在三个互补设置下评估视频分割能力:在MOSEv2数据集上进行复杂半监督视频对象分割,在MeViSv2-Text数据集上进行文本引导的指代视频对象分割,在MeViSv2-Audio数据集上进行音频引导的指代视频对象分割。报告描述了各项任务与评估协议,并回顾了各赛道前三名团队的方法。九个领先方案普遍结合基础分割模型与目标感知记忆、多模态推理、显式目标存在性验证、代理交互及修正跟踪机制。这些系统展现了从单一模型掩码传播向模块化流程的广泛转变,强调对物体身份、查询有效性与时间可靠性的综合推理。
原文摘要 · Abstract (English)
This report summarizes the 8th Large-scale Video Object Segmentation (LSVOS) Challenge, held in conjunction with ECCV 2026. The challenge evaluates video segmentation in three complementary settings: complex semi-supervised video object segmentation on MOSEv2, text-guided referring video object segmentation on MeViSv2-Text, and audio-guided referring video object segmentation on MeViSv2-Audio. We describe the tasks and evaluation protocols and review the methods of the top three teams in each track. Across the nine leading solutions, foundation segmentation models are combined with target-aware memory, multimodal reasoning, explicit target-existence verification, agentic interaction, and corrective tracking. These systems illustrate a broader transition from single-model mask propagation toward modular pipelines that reason about object identity, query validity, and temporal reliability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。