自动优化提示,让SAM2更好识别伪装物体
CamoSAM2: SAM2-oriented Prompt Auto-Refinement for Video Camouflaged Object Detection
- 融合运动与外观特征自动生成初始提示
- 多阶段优化提升掩码精度,mIoU提升8.0%~10.1%
- 适合视频伪装目标检测任务,推理速度快
Segment Anything Model 2 (SAM2) 作为基于提示的视频基础模型,在视频目标分割中表现优异。由于伪装物体与其背景高度相似,即使人类也难以分辨,导致在真实场景中使用 SAM2 进行自动化分割面临伪装感知和可靠提示生成的挑战。为此,我们提出 CamoSAM2,一种面向 SAM2 的运动-外观提示诱导与优化框架(MAPI)。首先,设计一个同时融合运动与外观线索的提示诱导器,实现更准确的初始预测。随后,提出针对 SAM2 的视频自适应多提示优化策略(AMPR),用于修正初始粗掩码中的提示误差,进一步生成优质提示。具体包括三步:伪装物体判定、关键提示帧选择和多提示构建。在两个基准数据集上的大量实验表明,CamoSAM2 显著优于现有方法,mIoU 提升 8.0% 和 10.1%。此外,该方法在同类 VCOD 模型中推理速度最快。
原文摘要 · Abstract (English)
The Segment Anything Model 2 (SAM2), a prompt-guided video foundation model, has remarkably performed in video object segmentation, drawing significant attention in the community. Due to the high similarity between camouflaged objects and their surroundings, which makes them difficult to distinguish even by the human eye, the application of SAM2 for automated segmentation in real-world scenarios faces challenges in camouflage perception and reliable prompts generation. To address these issues, we propose CamoSAM2, a motion-appearance prompt inducer (MAPI) and refinement framework to automatically generate and refine prompts for SAM2, enabling high-quality automatic detection and segmentation in VCOD task. Initially, we introduce a prompt inducer that simultaneously integrates motion and appearance cues to detect camouflaged objects, delivering more accurate initial predictions than existing methods. Subsequently, we propose a video-based adaptive multi-prompts refinement (AMPR) strategy tailored for SAM2, aimed at mitigating prompt error in initial coarse masks and further producing good prompts. Specifically, we introduce a novel three-step process to generate reliable prompts by camouflaged object determination, pivotal prompt frame selection, and multi-prompts formation. Extensive experiments conducted on two benchmark datasets demonstrate that our proposed model, CamoSAM2, significantly outperforms existing state-of-the-art methods, achieving increases of 8.0% and 10.1% in mIoU metric. Additionally, our method achieves the fastest inference speed compared to current VCOD models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。