arXiv:2607.16324cs.CVcs.LG2026-07中稿 · publication in the…

让疟原虫识别模型说出判断依据,用显微镜下的形态特征解释结果。

SGMCE: Segment-Grounded Morphological Concept Explanation for Malaria Parasite Species Identification in Thick Blood Smears

论文配图:SGMCE: Segment-Grounded Morphological Concept Explanation for Malaria Parasite Species Identification in Thick Blood Smears
图 1 · 摘自论文原文
  • 基于检测框裁剪图像,提取形状、颜色、核质等14项形态特征
  • 通过GPT-4o生成自然语言解释,准确率高达91%以上
  • 无需额外标注,适合临床医生验证AI诊断的可信度

疟疾诊断依赖于厚血涂片中疟原虫物种的精确识别,但深度学习检测器仅输出分类结果,缺乏形态学依据,限制了显微镜医师对单个病例的审核能力。本文提出SGMCE(段落锚定形态概念解释)框架,无需额外训练、无需形态标注或带标签的解释数据,即可生成基于厚血涂片形态的逐检测自然语言解释。对每个检测目标,提取掩码引导的图像缩略图,采用自适应阈值计算14项手工设计的计算机视觉形态特征(形状、颜色、核质、血红素色素),并结合世界卫生组织标准参考手册构建的特定知识库,向GPT-4o提问,生成结构化解释,说明支持当前物种的形态特征及排除其他物种的原因。通过四项自动评估指标验证:知识库一致性(KBC)、CV主张忠实度(CCF)、区分性得分(DS)、LLM作为裁判(LLMj)。采用物种感知的否定过滤语义评分规则解决临床描述与知识库术语之间的词汇不匹配问题。在139张厚血涂片、737个检测实例中,涵盖四种疟原虫物种及白细胞,平均KBC为0.91,平均DS为0.99,平均CCF为0.97,规则级CCF分析表明视觉-语言模型提出的主张与其引用的测量值高度一致。

原文摘要 · Abstract (English)

Malaria diagnosis in endemic regions depends on species-level identification of Plasmodium parasites in thick blood smears, but deep learning detectors classify detections without providing morphological evidence for their predictions, limiting the ability of microscopists to audit those predictions at the case level. We present SGMCE (Segment-Grounded Morphological Concept Explanation), a post-hoc explanation framework that requires no additional training, no morphological annotations, and no labelled explanation data, yet produces per-detection natural-language explanations anchored in thick-smear morphology. For each detection, SGMCE extracts mask-guided crop thumbnails, computes fourteen handcrafted computer-vision morphological features (shape, colour, chromatin, haemozoin pigment) using adaptive within-mask thresholds, and queries GPT-4o with both visual evidence and computed measurements, conditioned on a thick-smear-specific knowledge base compiled from the World Health Organization bench aids. The primary output is a structured explanation identifying which morphological features support the detected species and why the competing species are excluded. Explanations are validated by four automatic metrics: Knowledge-Base Consistency (KBC), CV-Claim Faithfulness (CCF), Discriminativeness Score (DS), and LLM-as-Judge (LLMj). A sentence-level semantic scoring rule with species-aware negation filtering resolves the vocabulary mismatch between clinical prose and knowledge-base terms. Across 737 detections from 139 thick-smear images spanning four Plasmodium species and white blood cells, parasite-class mean KBC is 0.91, mean DS is 0.99, and mean CCF is 0.97, while a per-rule CCF breakdown confirms that the CV-grounded claims made by the vision-language model are consistent with the measurements they cite.

医学AI可解释性疟疾诊断视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。