arXiv:2507.22792cs.CV2025-07综述被引 3

基于SAM模型的视频目标分割与追踪综述,梳理过去、现在与未来技术演进。

Segment Anything for Video: A Comprehensive Review of Video Object Segmentation and Tracking from Past to Future

  • 从历史信息保留到实时流式记忆,构建三阶段时间框架
  • 提出运动感知记忆选择与轨迹引导提示,提升精度与效率
  • 适合关注基础模型在视频理解中应用的研究者参考

视频目标分割与追踪(VOST)是计算机视觉中的关键挑战,需在时序动态帧间实现鲁棒的分割与追踪。传统方法在领域泛化性、时间一致性及计算效率方面存在不足。以分割一切模型(SAM)及其后续版本SAM2为代表的基础模型带来了范式转变,支持提示驱动的分割并具备强泛化能力。本文综述了基于SAM/SAM2的VOST方法,按时间维度分为过去、现在与未来三部分:回顾历史信息保持与更新策略(过去),分析当前帧中提取与优化判别特征的方法(现在),探讨未来帧中运动预测与轨迹估计机制(未来)。文中揭示了从早期记忆架构到SAM2流式记忆与实时分割能力的演进,并讨论了运动感知记忆选择与轨迹引导提示等新进展,旨在提升准确率与效率。最后指出仍存在的挑战,如记忆冗余、误差累积与提示低效,并展望未来研究方向。该综述为基于基础模型推动VOST发展的研究者提供了系统性指引。

原文摘要 · Abstract (English)

Video Object Segmentation and Tracking (VOST) presents a complex yet critical challenge in computer vision, requiring robust integration of segmentation and tracking across temporally dynamic frames. Traditional methods have struggled with domain generalization, temporal consistency, and computational efficiency. The emergence of foundation models like the Segment Anything Model (SAM) and its successor, SAM2, has introduced a paradigm shift, enabling prompt-driven segmentation with strong generalization capabilities. Building upon these advances, this survey provides a comprehensive review of SAM/SAM2-based methods for VOST, structured along three temporal dimensions: past, present, and future. We examine strategies for retaining and updating historical information (past), approaches for extracting and optimizing discriminative features from the current frame (present), and motion prediction and trajectory estimation mechanisms for anticipating object dynamics in subsequent frames (future). In doing so, we highlight the evolution from early memory-based architectures to the streaming memory and real-time segmentation capabilities of SAM2. We also discuss recent innovations such as motion-aware memory selection and trajectory-guided prompting, which aim to enhance both accuracy and efficiency. Finally, we identify remaining challenges including memory redundancy, error accumulation, and prompt inefficiency, and suggest promising directions for future research. This survey offers a timely and structured overview of the field, aiming to guide researchers and practitioners in advancing the state of VOST through the lens of foundation models.

视频分割目标追踪SAM模型基础模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。