用视觉语言模型让白血病细胞分析可解释,提升医生信任度
HemBLIP: A Vision-Language Model for Interpretable Leukemia Cell Morphology Analysis
- 基于1.4万张血细胞图像与专家标注,训练可生成描述的视觉语言模型
- 在形态准确性和描述质量上优于MedGEMMA,LoRA微调效率更高
- 适合临床辅助诊断、可解释性要求高的医学视觉任务
外周血细胞形态学的显微评估是白血病诊断的核心,但当前深度学习模型常为黑箱,限制了临床信任与应用。本文提出HemBLIP,一种用于生成可解释、形态感知的血细胞描述的视觉语言模型。基于包含14,000张健康与白血病细胞图像及专家属性标注的新数据集,通过全量微调和基于LoRA的参数高效训练方法适配通用视觉语言模型,并与生物医学基础模型MedGEMMA进行对比。HemBLIP在描述质量和形态准确性上表现更优,且LoRA微调在显著降低计算成本的同时带来额外性能提升。结果表明,视觉语言模型在透明化、可扩展的血液学诊断中具有巨大潜力。
原文摘要 · Abstract (English)
Microscopic evaluation of white blood cell morphology is central to leukemia diagnosis, yet current deep learning models often act as black boxes, limiting clinical trust and adoption. We introduce HemBLIP, a vision language model designed to generate interpretable, morphology aware descriptions of peripheral blood cells. Using a newly constructed dataset of 14k healthy and leukemic cells paired with expert-derived attribute captions, we adapt a general-purpose VLM via both full fine-tuning and LoRA based parameter efficient training, and benchmark against the biomedical foundation model MedGEMMA. HemBLIP achieves higher caption quality and morphological accuracy, while LoRA adaptation provides further gains with significantly reduced computational cost. These results highlight the promise of vision language models for transparent and scalable hematological diagnostics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。