用视觉大模型自动分析电子显微镜中的碳纳米管形态,又快又准。
Quantification and Classification of Carbon Nanotubes in Electron Micrographs using Vision Foundation Models
- 基于SAM的交互式分割工具,少量标注即可精准提取颗粒。
- 结合DINOv2实现95.5%分类准确率,仅需少量训练数据。
- 可分辨同一视野中混合存在的多种碳纳米管类型,适合材料研究者。
碳纳米管形貌的精确表征对暴露评估和毒理学研究至关重要,但现有流程依赖耗时且主观的人工分割。本文提出一个统一框架,利用视觉基础模型实现电子显微镜图像中碳纳米管的自动化定量与分类。首先,基于分割一切模型(SAM)构建交互式量化工具,仅需极少用户输入即可实现近乎完美的颗粒分割。其次,提出新型分类流水线,利用分割掩码空间约束DINOv2视觉变换器,仅从粒子区域提取特征并抑制背景噪声。在1,800张透射电镜(TEM)图像数据集上,该架构在区分四种不同碳纳米管形貌任务中达到95.5%准确率,显著优于当前基线,且训练数据量仅为后者的极小部分。关键的是,该实例级处理能力可解析混合样本,正确识别单个视场中共存的不同粒子类型。结果表明,融合零样本分割与自监督特征学习,可实现高通量、可复现的纳米材料分析,将原本劳动密集型的瓶颈转变为可扩展的数据驱动流程。
原文摘要 · Abstract (English)
Accurate characterization of carbon nanotube morphologies in electron microscopy images is vital for exposure assessment and toxicological studies, yet current workflows rely on slow, subjective manual segmentation. This work presents a unified framework leveraging vision foundation models to automate the quantification and classification of CNTs in electron microscopy images. First, we introduce an interactive quantification tool built on the Segment Anything Model (SAM) that segments particles with near-perfect accuracy using minimal user input. Second, we propose a novel classification pipeline that utilizes these segmentation masks to spatially constrain a DINOv2 vision transformer, extracting features exclusively from particle regions while suppressing background noise. Evaluated on a dataset of 1,800 TEM images, this architecture achieves 95.5% accuracy in distinguishing between four different CNT morphologies, significantly outperforming the current baseline despite using a fraction of the training data. Crucially, this instance-level processing allows the framework to resolve mixed samples, correctly classifying distinct particle types co-existing within a single field of view. These results demonstrate that integrating zero-shot segmentation with self-supervised feature learning enables high-throughput, reproducible nanomaterial analysis, transforming a labor-intensive bottleneck into a scalable, data-driven process.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。