基于SeC框架实现长期与语义感知的视频分割,获2025年MOSEv2挑战赛第一名。
The 1st Solution for MOSEv2 Challenge 2025: Long-term and Concept-aware Video Segmentation via SeC
- 采用增强版SAM-2框架SeC,融合长期记忆与语义先验机制。
- 在测试集上取得39.89%的JF分数,显著优于其他方法。
- 适合关注长时序视频分割与抗干扰能力的研究者。
本技术报告探讨了LSVOS挑战赛中的MOSEv2赛道,该赛道聚焦复杂半监督视频对象分割任务。通过分析并改进SeC——一种增强版SAM-2框架,我们深入研究其长期记忆与概念感知记忆机制。结果表明,长期记忆在遮挡和重出现场景下保持时间连续性,而概念感知记忆则提供语义先验以抑制干扰项;二者协同有效应对多个MOSEv2核心挑战。我们的方案在测试集上达到39.89%的JF得分,位列MOSEv2赛道第一。
原文摘要 · Abstract (English)
This technical report explores the MOSEv2 track of the LSVOS Challenge, which targets complex semi-supervised video object segmentation. By analysing and adapting SeC, an enhanced SAM-2 framework, we conduct a detailed study of its long-term memory and concept-aware memory, showing that long-term memory preserves temporal continuity under occlusion and reappearance, while concept-aware memory supplies semantic priors that suppress distractors; together, these traits directly benefit several MOSEv2's core challenges. Our solution achieves a JF score of 39.89% on the test set, ranking 1st in the MOSEv2 track of the LSVOS Challenge.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。