SAM2在视频伪装物体分割任务中表现优异,微调后效果更佳。
When SAM2 Meets Video Camouflaged Object Segmentation: A Comprehensive Evaluation and Adaptation
- 用点击、框选和掩码提示评估SAM2在伪装视频中的表现
- 零样本下可有效识别伪装物体,微调后性能进一步提升
- 适合关注视频目标分割与模型适配的研究者
本研究系统评估了视频基础模型SAM2在视频伪装物体分割(VCOS)任务中的表现。VCOS需在颜色、纹理相似及光照不佳等条件下检测与背景融为一体的物体,难度远高于常规场景。本文首先在多个伪装视频数据集上测试不同提示方式(点击、框选、掩码)下的SAM2性能;其次探索其与多模态大语言模型及现有VCOS方法的融合潜力;最后针对视频伪装数据集对SAM2进行微调。实验表明,SAM2具备出色的零样本分割能力,且通过特定参数调整可进一步提升性能。代码已开源。
原文摘要 · Abstract (English)
This study investigates the application and performance of the Segment Anything Model 2 (SAM2) in the challenging task of video camouflaged object segmentation (VCOS). VCOS involves detecting objects that blend seamlessly in the surroundings for videos, due to similar colors and textures, poor light conditions, etc. Compared to the objects in normal scenes, camouflaged objects are much more difficult to detect. SAM2, a video foundation model, has shown potential in various tasks. But its effectiveness in dynamic camouflaged scenarios remains under-explored. This study presents a comprehensive study on SAM2's ability in VCOS. First, we assess SAM2's performance on camouflaged video datasets using different models and prompts (click, box, and mask). Second, we explore the integration of SAM2 with existing multimodal large language models (MLLMs) and VCOS methods. Third, we specifically adapt SAM2 by fine-tuning it on the video camouflaged dataset. Our comprehensive experiments demonstrate that SAM2 has excellent zero-shot ability of detecting camouflaged objects in videos. We also show that this ability could be further improved by specifically adjusting SAM2's parameters for VCOS. The code is available at https://github.com/zhoustan/SAM2-VCOS
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。