通过竞争性读取记忆,提升复杂视频中目标的分割鲁棒性。
Competitive Memory Readout for Robust Video Object Segmentation: 2nd Place Technical Report for the MOSEv2 Track of the 8th LSVOS Challenge

- 引入竞争性记忆读取,显式建模同类别干扰物影响
- 在长时遮挡等场景下实现66.20的主指标得分
- 适合处理目标外观变化大、同类干扰多的视频分割任务
我们提出针对ECCV 2026第八届大规模视频对象分割(LSVOS)挑战赛MOSEv2赛道的解决方案。该挑战评估在复杂时序动态下的鲁棒视频对象分割能力,包括长期遮挡、消失与重现、显著外观变化以及视觉相似物体的强干扰。我们的方法基于SAM~3,聚焦于其记忆读取机制。标准的目标单一记忆检索会因同类别非目标物体仅以背景隐式表示而产生混淆。为此,我们提出竞争性记忆读取(Competitive Memory Readout),在从记忆中检索目标信息时显式引入同类别竞争者证据。为防止对弱目标或重现目标过度抑制,我们在竞争后应用轻量级自适应恢复规则。该系统保留了原始SAM~3追踪流程,同时提升了复杂视频中目标身份的保持能力。我们的提交在主指标上取得66.20分,位列MOSEv2赛道第二名。
原文摘要 · Abstract (English)
We present our solution for the MOSEv2 track of the 8th Large-scale Video Object Segmentation (LSVOS) Challenge at ECCV 2026. The challenge evaluates robust video object segmentation under complex temporal dynamics, including long-term occlusion, disappearance and reappearance, large appearance changes, and strong interference from visually similar objects. Our method builds on SAM~3 and focuses on its memory readout. Standard target-only memory retrieval can confuse the annotated target with same-class non-target objects because such distractors are represented only implicitly as background. Our method introduces Competitive Memory Readout, which explicitly incorporates same-class competitor evidence when retrieving target information from memory. To prevent excessive suppression of weak or reappearing targets, we further apply a lightweight adaptive restoration rule after competition. The resulting system retains the original SAM~3 tracking pipeline while improving target identity preservation in challenging videos. Our submission achieves 66.20 on the primary challenge score and ranks 2nd in the MOSEv2 track.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。