用视觉大模型零样本分类纳米颗粒形状,无需标注数据。
Zero-shot Shape Classification of Nanoparticles in SEM Images using Vision Foundation Models
- 用SAM分割+DINOv2提取特征,搭配轻量分类器实现零样本识别。
- 在三个数据集上精度高,对微小形貌差异和领域迁移鲁棒。
- 适合缺乏标注数据的科研与工业用户快速分析纳米颗粒形态。
准确高效地表征扫描电子显微镜(SEM)图像中纳米颗粒的形貌,对保障纳米材料合成质量及加速研发至关重要。然而,传统深度学习方法需大量标注数据和耗时训练,限制了科研与工业界普通研究者使用。本研究提出一种零样本分类流程,利用两个视觉基础模型:用于目标分割的Segment Anything Model(SAM)和用于特征嵌入的DINOv2。结合轻量分类器,该方法在无需大规模参数微调的情况下,实现了对三类形貌多样纳米颗粒数据集的高精度形状分类。相比微调后的YOLOv11和ChatGPT o4-mini-high基线,本方法在小样本、细微形貌变化及自然图像到科学成像域转移下均表现更优。通过DINOv2特征的PCA聚类定量分析,可评估化学合成进程。本工作展示了基础模型在自动化显微图像分析中的潜力,为纳米颗粒研究提供了一种更高效、更易用的替代方案。
原文摘要 · Abstract (English)
Accurate and efficient characterization of nanoparticle morphology in Scanning Electron Microscopy (SEM) images is critical for ensuring product quality in nanomaterial synthesis and accelerating development. However, conventional deep learning methods for shape classification require extensive labeled datasets and computationally demanding training, limiting their accessibility to the typical nanoparticle practitioner in research and industrial settings. In this study, we introduce a zero-shot classification pipeline that leverages two vision foundation models: the Segment Anything Model (SAM) for object segmentation and DINOv2 for feature embedding. By combining these models with a lightweight classifier, we achieve high-precision shape classification across three morphologically diverse nanoparticle datasets - without the need for extensive parameter fine-tuning. Our methodology outperforms a fine-tuned YOLOv11 and ChatGPT o4-mini-high baselines, demonstrating robustness to small datasets, subtle morphological variations, and domain shifts from natural to scientific imaging. Quantitative clustering metrics on PCA plots of the DINOv2 features are discussed as a means of assessing the progress of the chemical synthesis. This work highlights the potential of foundation models to advance automated microscopy image analysis, offering an alternative to traditional deep learning pipelines in nanoparticle research which is both more efficient and more accessible to the user.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。