arXiv:2510.11260cs.CVcond-mat.mtrl-sci2025-10

用大模型自动识别扫描电镜图像中的比例尺,提升分析效率与准确性。

A Large-Language-Model Assisted Automated Scale Bar Detection and Extraction Framework for Scanning Electron Microscopic Images

  • 融合多模态检测与大模型推理,实现比例尺定位与文字识别一体化。
  • 比例尺检测精度达100%,召回率95.8%,在复杂条件下仍保持高准确率。
  • 适合科研人员快速处理大量电镜图像,尤其适合自动化分析需求者。

扫描电子显微镜(SEM)广泛用于微观结构的可视化与分析。确定比例尺是准确分析的重要前提,但当前主要依赖人工操作,耗时且易出错。为此,我们提出一种基于大语言模型(LLM)的多模态自动化比例尺检测与提取框架,包含四个阶段:(i)自动数据生成(Auto-DG)模型合成多样化SEM图像以增强训练鲁棒性;(ii)比例尺目标检测;(iii)结合DenseNet与卷积循环神经网络(CRNN)的混合光学字符识别(OCR)系统进行信息提取;(iv)LLM代理对结果进行分析与验证。该框架在目标检测中达到精确率100%、召回率95.8%、平均精度均值(mAP)99.2%(IoU=0.5)和69.1%(IoU=0.5:0.95)。混合OCR系统在Auto-DG数据集上实现89%精确率、65%召回率与75%F1分数,显著优于多个主流独立引擎。LLM作为推理引擎与智能助手,可建议后续步骤并验证结果。该方法显著提升比例尺检测与提取的效率与准确性,为科学成像分析提供有力工具。

原文摘要 · Abstract (English)

Microscopic characterizations, such as Scanning Electron Microscopy (SEM), are widely used in scientific research for visualizing and analyzing microstructures. Determining the scale bars is an important first step of accurate SEM analysis; however, currently, it mainly relies on manual operations, which is both time-consuming and prone to errors. To address this issue, we propose a multi-modal and automated scale bar detection and extraction framework that provides concurrent object detection, text detection and text recognition with a Large Language Model (LLM) agent. The proposed framework operates in four phases; i) Automatic Dataset Generation (Auto-DG) model to synthesize a diverse dataset of SEM images ensuring robust training and high generalizability of the model, ii) scale bar object detection, iii) information extraction using a hybrid Optical Character Recognition (OCR) system with DenseNet and Convolutional Recurrent Neural Network (CRNN) based algorithms, iv) an LLM agent to analyze and verify accuracy of the results. The proposed model demonstrates a strong performance in object detection and accurate localization with a precision of 100%, recall of 95.8%, and a mean Average Precision (mAP) of 99.2% at IoU=0.5 and 69.1% at IoU=0.5:0.95. The hybrid OCR system achieved 89% precision, 65% recall, and a 75% F1 score on the Auto-DG dataset, significantly outperforming several mainstream standalone engines, highlighting its reliability for scientific image analysis. The LLM is introduced as a reasoning engine as well as an intelligent assistant that suggests follow-up steps and verifies the results. This automated method powered by an LLM agent significantly enhances the efficiency and accuracy of scale bar detection and extraction in SEM images, providing a valuable tool for microscopic analysis and advancing the field of scientific imaging.

图像识别大模型应用科学成像自动化分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。