不训练模型,用双记忆机制提升长视频分割稳定性
SAM2Dual: Training-Free, Dual Memory for Long-Term Video Object Segmentation

- 采用短时与长时双记忆分离设计,分别应对局部变化和全局身份保持
- 在MOSEv2上将J&F从49.33提升至50.65,显著降低长期遮挡下的误差累积
- 适合需要高鲁棒性的长视频分割场景,尤其视觉线索弱时仍能保持身份一致
长时视频目标分割(VOS)因长期遮挡、物体重现和场景变化面临挑战。尽管SAM2具备强零样本性能,但其流式记忆在长时间序列中易受近期不可靠预测影响,导致漂移加剧。我们提出SAM2Dual,一种无需训练、可即插即用的推理时增强方法,不修改模型权重即可提升长视频鲁棒性。SAM2Dual采用双记忆设计:短时记忆用于快速局部适应,长时记忆通过间隔采样构建以保留全局身份线索,并通过门控融合策略结合两者。此外,提出文本感知记忆(TAM),从早期帧提取词级提示,利用文本嵌入按语义兼容性重加权记忆贡献,在视觉证据弱或模糊时支持身份保持。在多个长时基准测试中,SAM2Dual持续提升稳定性,使MOSEv2上的J&F从49.33提升至50.65,并在LVOSv2上实现一致增益。
原文摘要 · Abstract (English)
Long-term video object segmentation (VOS) remains challenging due to error accumulation under extended occlusions, re-appearance, and scene changes. Although SAM2 provides strong zero-shot performance, its streaming memory can amplify drift over long horizons when recent, unreliable predictions dominate the memory state. We propose SAM2Dual, a training-free, plug-and-play inference-time enhancement that improves long-video robustness without updating model weights. SAM2Dual introduces a Dual Memory design that explicitly separates (i) short-term memory for rapid local adaptation and (ii) long-term memory built via interval-based sampling to preserve global identity cues, combined through a gated fusion strategy. In addition, we present Text-Aware Memory (TAM), which extracts a compact word-level cue from early frames and uses text embeddings to reweight memory contributions based on semantic compatibility, supporting identity preservation when visual evidence becomes weak or ambiguous. Across long-term benchmarks, SAM2Dual consistently improves stability on long videos, raising J&F from 49.33 to 50.65 on MOSEv2 and achieving consistent gains on LVOSv2.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。