用图像模型实现视频抠图,兼顾精度与连贯性。
SAM2Matting: Generalized Image and Video Matting

- 用追踪器+区域提议桥接,分离追踪与抠图任务
- 仅用图像训练却达视频抠图新基准
- 适合需要高精度、强泛化的视频处理场景
尽管图像抠图已取得显著进展,视频抠图仍因高层追踪(需逐帧理解)与低层抠图(关注细微细节)之间的固有差距而面临挑战。现有方法依赖昂贵且范围受限的视频抠图数据集,可能限制域外泛化能力并影响追踪鲁棒性。我们提出SAM2Matting,一种从追踪到抠图的新型范式,将VOS追踪器升级为高保真视频抠图模型。具体而言,该框架通过增强基础追踪器(如SAM2、SAM3),引入区域提议桥接和专用抠图头,使追踪器专注时间一致性,抠图组件负责精细细节。值得注意的是,尽管仅在图像上训练,SAM2Matting在视频抠图任务上达到新最优性能,支持多种提示类型,保持强时间连贯性,并在以人物为中心及野外场景中均展现良好泛化能力。
原文摘要 · Abstract (English)
Despite impressive advances in image matting, video matting remains challenging due to the inherent gap between high-level tracking, which requires frame-wise understanding, and low-level matting, which focuses on extremely fine-grained details. Existing methods attempt this with expensive and narrowly-scoped video matting datasets, which may limit out-of-domain generalization and compromise tracking robustness. We rethink the paradigm with SAM2Matting, a tracker-to-matting framework that advances VOS trackers to high-fidelity video matting. Specifically, it decouples the task by enhancing a foundational tracker (e.g., SAM2, SAM3) with a region-proposal bridge and dedicated matting heads, enabling the uncompromised tracker to handle temporal consistency while the matting components resolve fine-grained details. Notably, despite being trained only on images, SAM2Matting establishes new state-of-the-art performance on video matting, supports diverse prompt types, maintains strong temporal consistency, and demonstrates robust generalization across both human-centric and in-the-wild scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。