基于SAM的统一框架,实现物体完整形状的泛化分割。
Amodal SAM: A Unified Amodal Segmentation Framework with Generalization

- 轻量级空间补全适配器重建遮挡区域。
- 合成数据生成解决标注稀缺问题,提升泛化能力。
- 新学习目标保证区域一致性和拓扑正确性。
非可见分割旨在预测物体完整几何形状,包括被遮挡部分。现有方法多局限于训练域内,难以泛化至新类别和未见场景。本文提出Amodal SAM,一个基于SAM的统一框架,用于图像与视频的非可见分割。该框架在保持SAM强大泛化能力的同时,扩展其功能以支持非可见分割。改进体现在三方面:(1) 轻量级空间补全适配器,实现遮挡区域重建;(2) 目标感知遮挡合成(TAOS)流程,通过生成多样合成数据缓解非可见标注稀缺;(3) 新学习目标,强化区域一致性与拓扑正则化。大量实验表明,Amodal SAM在标准基准上达到最先进性能,且对新场景具有鲁棒泛化能力。本研究有望推动非可见分割向真实世界应用迈进。
原文摘要 · Abstract (English)
Amodal segmentation is a challenging task that aims to predict the complete geometric shape of objects, including their occluded regions. Although existing methods primarily focus on amodal segmentation within the training domain, these approaches often lack the generalization capacity to extend effectively to novel object categories and unseen contexts. This paper introduces Amodal SAM, a unified framework that leverages SAM (Segment Anything Model) for both amodal image and amodal video segmentation. Amodal SAM preserves the powerful generalization ability of SAM while extending its inherent capabilities to the amodal segmentation task. The improvements lie in three aspects: (1) a lightweight Spatial Completion Adapter that enables occluded region reconstruction, (2) a Target-Aware Occlusion Synthesis (TAOS) pipeline that addresses the scarcity of amodal annotations by generating diverse synthetic training data, and (3) novel learning objectives that enforce regional consistency and topological regularization. Extensive experiments demonstrate that Amodal SAM achieves state-of-the-art performance on standard benchmarks, while simultaneously exhibiting robust generalization to novel scenarios. We anticipate that this research will advance the field toward practical amodal segmentation systems capable of operating effectively in unconstrained real-world environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。