arXiv:2608.00361cs.CVcs.AI2026-08被引 1

用AI从显微图像中提取可解释的颗粒物语义特征,提升分类与检索精度。

Artificial Intelligence for the Characterization of Particles and Fibers by Optical Microscopy

  • 通过多模态教师模型融合图像与光照、放大倍数等文本信息生成可解释嵌入
  • 学生模型在仅输入图像时达到80%伪类别准确率和75%召回率
  • 无需对比学习即可防止特征坍缩,适合复杂颗粒物的分析与检索

颗粒与纤维分散体的光学显微观察依赖于受样品形貌、化学成分、放大倍数和照明条件影响的细微视觉线索。我们提出一种人工智能蒸馏框架,利用语义锚点从显微图像中提取语义丰富的图像嵌入。一个多模态教师模型将每张图像的视觉嵌入与代表照明模式、放大倍数以及样品身份和形貌的三个文本嵌入结合,由LongCLIP扩展上下文文本编码器生成一个2304维的块结构教师向量,其组件块在训练和推理过程中保持物理可解释性。一个带有多层感知机解码器的学生视觉变换器(ViT)被训练仅从图像重建该教师向量,最小化均绝对误差(L1)损失以确保坐标级保真度。对教师嵌入空间进行HDBSCAN聚类得到伪类别,通过交叉熵项作为防坍缩正则化,无需对比负样本挖掘即实现类间分离。推理时,学生模型仅需图像输入,生成紧凑嵌入并恢复教师向量的完整语义内容。该框架在留一法最近邻检索中实现了约80%的伪类别验证准确率和75%的Recall@1,证明语义锚定使仅视觉学生模型获得比纯图像训练更丰富且更可解释的表征,适用于异质颗粒与纤维分散体的检索、分类及探索性分析。

原文摘要 · Abstract (English)

Optical microscopy of particle and fiber dispersions involves interpreting subtle visual cues influenced by specimen morphology, chemical composition, magnification, and illumination conditions. We introduce an artificial intelligence (AI) distillation framework that extracts semantically rich image embeddings from microscopy images using semantic anchors. A multimodal teacher combines each image's visual embedding with three text embeddings representing illumination modality, magnification, and specimen identity and morphology. Generated by LongCLIP's extended-context text encoder, this yields a 2304-dimensional block-structured teacher vector whose component blocks remain physically interpretable throughout training and inference. A student vision transformer (ViT) with a multi-layer perceptron (MLP) decoder is trained to reconstruct this teacher vector from the image alone, minimizing a mean absolute error (L1) loss that enforces coordinate-level fidelity to the teacher's block structure. A cross-entropy term over pseudo-classes derived from HDBSCAN clustering of the teacher embedding space acts as a collapse-prevention regularizer, enforcing inter-cluster separation without requiring contrastive negative mining. At inference, the student operates on image input alone, producing compact embeddings that recover the full semantic content of the teacher vector. The framework achieves approximately 80% pseudo-class validation accuracy and 75% Recall@1 on fine-grained specimen description labels under leave-one-out nearest-neighbor retrieval. These results demonstrate that semantic anchoring enables a vision-only student to acquire richer and more interpretable representations than image-only training, with direct applicability to retrieval, classification, and exploratory analysis of heterogeneous particle and fiber dispersions.

AI显微颗粒分析语义嵌入蒸馏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。