用少量标注样本实现机械部件精准分割,支持真实场景泛化。
Few-shot Structure-Informed Machinery Part Segmentation with Foundation Models and Graph Neural Networks
- 结合视觉基础模型与图网络,利用空间层级关系提升分割精度。
- 仅需1-25个标注样本,合成数据上分割效果优秀,真实数据上J&F达92.2。
- 训练快(<5分钟),适合自动驾驶等需快速部署的机械交互系统。
本文提出一种新颖的少样本语义分割方法,用于具有空间与层级关系的多部件机械装置。通过融合CLIPSeg、Segment Anything Model(SAM)和SuperPoint兴趣点检测器,结合图卷积网络(GCN),实现对机械部件的精准分割。在纯合成的卡车式装载起重机数据集上,仅需1至25个标注样本即可实现多级细节的高效分割。模型在消费级GPU上训练时间低于五分钟。在真实数据上表现出强泛化能力,使用10个合成支持样本时,$J\&F$得分达到92.2。在DAVIS 2017数据集上,半监督视频分割任务中使用三个支持样本,$J\&F$得分为71.5。该方法训练迅速且能有效迁移到真实场景,适用于与机械和基础设施交互的自主系统,展示了组合式基础模型在少样本分割中的潜力。
原文摘要 · Abstract (English)
This paper proposes a novel approach to few-shot semantic segmentation for machinery with multiple parts that exhibit spatial and hierarchical relationships. Our method integrates the foundation models CLIPSeg and Segment Anything Model (SAM) with the interest point detector SuperPoint and a graph convolutional network (GCN) to accurately segment machinery parts. By providing 1 to 25 annotated samples, our model, evaluated on a purely synthetic dataset depicting a truck-mounted loading crane, achieves effective segmentation across various levels of detail. Training times are kept under five minutes on consumer GPUs. The model demonstrates robust generalization to real data, achieving a qualitative synthetic-to-real generalization with a $J\&F$ score of 92.2 on real data using 10 synthetic support samples. When benchmarked on the DAVIS 2017 dataset, it achieves a $J\&F$ score of 71.5 in semi-supervised video segmentation with three support samples. This method's fast training times and effective generalization to real data make it a valuable tool for autonomous systems interacting with machinery and infrastructure, and illustrate the potential of combined and orchestrated foundation models for few-shot segmentation tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。