用强化学习优化SAM2的记忆更新,提升视频追踪精度。
SAM2RL: Towards Reinforcement Learning Memory Control in Segment Anything Model 2
- 将记忆控制建模为序列决策问题,用强化学习自动优化
- 在过拟合设置下,性能提升超现有方法三倍以上
- 适合需要高精度追踪的视觉任务研究者
Segment Anything Model 2(SAM 2)在物体分割任务中表现出色,已成为视觉目标跟踪的最先进模型。该模型通过记忆库存储前帧信息,实现视频序列中的时序一致性。近期方法通过人工设计的更新规则来应对干扰、遮挡和运动变化。本文提出一种根本性不同的方法:将记忆控制视为序列决策问题,利用强化学习优化记忆更新。在每个视频配备独立智能体的过拟合设置下,本方法相对于SAM 2的性能提升超过现有启发式方法提升的三倍。结果揭示了记忆库的未开发潜力,并表明强化学习是视觉目标跟踪中记忆控制的有力替代方案。
原文摘要 · Abstract (English)
Segment Anything Model 2 (SAM 2) has demonstrated strong performance in object segmentation tasks and has become the state-of-the-art for visual object tracking. The model stores information from previous frames in a memory bank, enabling temporal consistency across video sequences. Recent methods augment SAM 2 with hand-crafted update rules to better handle distractors, occlusions, and object motion. We propose a fundamentally different approach using reinforcement learning for optimizing memory updates in SAM 2 by framing memory control as a sequential decision-making problem. In an overfitting setup with a separate agent per video, our method achieves a relative improvement over SAM 2 that exceeds by more than three times the gains of existing heuristics. These results reveal the untapped potential of the memory bank and highlight reinforcement learning as a powerful alternative to hand-crafted update rules for memory control in visual object tracking.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。