无需训练,单张示例即可精准分割物体细部
One-shot In-context Part Segmentation
- 利用DINOv2与Stable Diffusion互补特征,自适应选择通道增强细节区分力
- 在三个基准数据集上超越现有方法,对视角和遮挡变化具有强鲁棒性
- 适合少样本、快速部署的细粒度分割场景
本文提出One-shot In-context Part Segmentation(OIParts)框架,解决基于视觉基础模型(VFMs)的一次性细部分割在外观、视角差异大或物体部分遮挡时的泛化难题。传统训练式方法易过拟合,限制泛化能力。OIParts采用无训练、数据高效的设计,仅需一个上下文示例即可实现高精度分割。通过融合DINOv2与Stable Diffusion的互补特征,提出最小化类内距离的自适应通道选择策略,提升细粒度特征判别力。在三个基准数据集上的实验证明,该方法显著优于现有一次分割方法,具备更强泛化能力,且无需大量标注数据。
原文摘要 · Abstract (English)
In this paper, we present the One-shot In-context Part Segmentation (OIParts) framework, designed to tackle the challenges of part segmentation by leveraging visual foundation models (VFMs). Existing training-based one-shot part segmentation methods that utilize VFMs encounter difficulties when faced with scenarios where the one-shot image and test image exhibit significant variance in appearance and perspective, or when the object in the test image is partially visible. We argue that training on the one-shot example often leads to overfitting, thereby compromising the model's generalization capability. Our framework offers a novel approach to part segmentation that is training-free, flexible, and data-efficient, requiring only a single in-context example for precise segmentation with superior generalization ability. By thoroughly exploring the complementary strengths of VFMs, specifically DINOv2 and Stable Diffusion, we introduce an adaptive channel selection approach by minimizing the intra-class distance for better exploiting these two features, thereby enhancing the discriminatory power of the extracted features for the fine-grained parts. We have achieved remarkable segmentation performance across diverse object categories. The OIParts framework not only eliminates the need for extensive labeled data but also demonstrates superior generalization ability. Through comprehensive experimentation on three benchmark datasets, we have demonstrated the superiority of our proposed method over existing part segmentation approaches in one-shot settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。