改进视觉相似物体追踪,提升长视频分割稳定性
SAMSON: 3rd Place Solution of LSVOS 2025 VOS Challenge
- 引入长效记忆模块,解决物体重识别难题
- 采用SAM2Long后处理,降低误差累积,提升分割精度
- 在MOSE挑战赛中获第三名,测试集J&F达0.8427
大规模视频目标分割(LSVOS)旨在精准追踪和分割长视频序列中的物体,难点包括物体重现、小目标、严重遮挡和密集场景。现有方法多基于SAM2框架并结合各类记忆机制进行复杂视频掩码生成。本文提出一种名为「带记忆强化目标导航的通用分割模型」(SAMSON)的解决方案,为ICCV 2025 LSVOS MOSE赛道第三名。该方法融合先进VOS模型优势,构建有效范式:通过长效记忆模块实现对视觉相似实例与长期消失物体的可靠重识别;同时采用SAM2Long作为后处理策略,减少长视频中的误差累积,增强分割稳定性。最终在测试集排行榜上取得J&F为0.8427的性能表现。
原文摘要 · Abstract (English)
Large-scale Video Object Segmentation (LSVOS) addresses the challenge of accurately tracking and segmenting objects in long video sequences, where difficulties stem from object reappearance, small-scale targets, heavy occlusions, and crowded scenes. Existing approaches predominantly adopt SAM2-based frameworks with various memory mechanisms for complex video mask generation. In this report, we proposed Segment Anything with Memory Strengthened Object Navigation (SAMSON), the 3rd place solution in the MOSE track of ICCV 2025, which integrates the strengths of stateof-the-art VOS models into an effective paradigm. To handle visually similar instances and long-term object disappearance in MOSE, we incorporate a long-term memorymodule for reliable object re-identification. Additionly, we adopt SAM2Long as a post-processing strategy to reduce error accumulation and enhance segmentation stability in long video sequences. Our method achieved a final performance of 0.8427 in terms of J &F in the test-set leaderboard.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。