微调SAM模型可显著提升垃圾分类分割效果
Don't waste SAM

- 用三类真实垃圾数据微调SAM,增强泛化能力
- 微调后SAM在三个数据集上IoU提升30点,接近顶尖模型
- 适合需要快速构建分割系统的开发者和研究者
Meta AI发布的通用图像分割模型SAM在零样本场景下表现出色,但在实际应用中仍面临遮挡、形变、透明及背景混淆等挑战。本研究评估了SAM在三个真实场景垃圾数据集上的表现,并通过微调其ViT-H版本进行优化。结果显示,微调后的SAM在Zerowaste和TACO数据集上平均IoU提升30点,性能接近TrashCan 1.0,仅差-1.44。实验表明,将SAM作为基础模型进行微调,是实现更好下游任务泛化的关键步骤。因此,不应忽视或浪费SAM的价值。
原文摘要 · Abstract (English)
Meta AI has recently released the Segment Anything Model (SAM), which demonstrates exceptional zero-shot image segmentation performance across various tasks with remarkable accuracy. Despite its inability to provide accurate segmentation across multiple research fields, SAM still serves as a valuable starting point for supporting the segmentation pipeline process, particularly for tasks that require extensive and senior skills annotations. This study aims to evaluate the generalization of SAM and fine-tuning SAM models using three waste segmentation datasets. Although they are captured from real scenes as SAM was pretrained on, these datasets present several challenges, including occlusions, deformable objects, transparency, and objects easily confused with backgrounds. In our findings, the fine-tuned SAM-ViT-H model outperforms the state-ofthe-art Zerowaste, and TACO datasets with a significant increase of +30 in IoU, and it closely approaches performance levels of TrashCan 1.0, with only a -1.44 difference. After evaluating these popular waste datasets, it became evident that fine-tuning SAM as a foundational model is a crucial step for providing better generalization for downstream waste segmentation tasks. Therefore, SAM should not be disregarded or wasted.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。