arXiv:2602.13712cs.CVcs.LG2026-02

用微调的视觉语言模型自动定位显微镜下的寄生虫卵,提升诊断效率。

Fine-tuned Vision Language Model for Localization of Parasitic Eggs in Microscopic Images

  • 微调微软Florence模型实现寄生虫卵定位
  • mIOU达0.94,优于EfficientDet等检测方法
  • 适合资源有限地区的智能寄生虫病诊断应用

土壤传播线虫感染持续影响全球大量人口,尤其在热带和亚热带地区,当地缺乏专业诊断人才。尽管显微镜下人工检测寄生虫卵仍是诊断金标准,但该方法耗时费力且易出错。本文旨在利用微软Florence等视觉语言模型(VLM),通过微调实现对显微图像中所有寄生虫卵的定位。初步结果表明,该定位VLM性能优于EfficientDet等目标检测方法,mIOU达0.94。这一成果展示了所提VLM作为自动化诊断框架核心组件的潜力,为智能寄生虫学诊断提供可扩展的工程解决方案。

原文摘要 · Abstract (English)

Soil-transmitted helminth (STH) infections continuously affect a large proportion of the global population, particularly in tropical and sub-tropical regions, where access to specialized diagnostic expertise is limited. Although manual microscopic diagnosis of parasitic eggs remains the diagnostic gold standard, the approach can be labour-intensive, time-consuming, and prone to human error. This paper aims to utilize a vision language model (VLM) such as Microsoft Florence that was fine-tuned to localize all parasitic eggs within microscopic images. The preliminary results show that our localization VLM performs comparatively better than the other object detection methods, such as EfficientDet, with an mIOU of 0.94. This finding demonstrates the potential of the proposed VLM to serve as a core component of an automated framework, offering a scalable engineering solution for intelligent parasitological diagnosis.

视觉语言模型医学图像寄生虫检测自动化诊断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。