将医学专家知识融入视觉语言模型,提升先天性巨结肠病理检测精度。
Knowledge-Driven Vision-Language Model for Plexus Detection in Hirschsprung's Disease
- 用专家文本提示引导对比学习的视觉语言模型,融合临床语义与图像特征。
- 在病理分类中达到83.9%准确率、86.6%精确率、87.6%特异度,优于传统CNN模型。
- 适合病理医生辅助诊断,推动可解释性医学AI发展。
先天性巨结肠症是由于结肠部分区域神经节细胞先天缺失所致,导致该段肠道无法协调蠕动排便,常引发梗阻。其诊断与治疗需在组织切片显微视图中清晰识别肌间神经丛区域(myenteric plexus)是否含有神经节细胞。尽管卷积神经网络等深度学习方法在此任务中表现优异,但常被视为黑箱,缺乏可解释性,且不完全符合医生决策逻辑。本研究提出一种新框架,将专家提取的文本概念融入基于对比语言-图像预训练的视觉语言模型,以指导神经丛分类。通过大语言模型生成并经团队审核的专家来源提示(如医学教材和论文),经QuiltNet编码后对齐临床相关语义与视觉特征。实验结果表明,该模型在多项分类指标上均优于基于CNN的模型(包括VGG-19、ResNet-18和ResNet-50),实现83.9%的准确率、86.6%的精确率和87.6%的特异性。研究凸显了多模态学习在组织病理学中的潜力,并强调融入专家知识对生成更具临床意义模型输出的重要性。
原文摘要 · Abstract (English)
Hirschsprung's disease is defined as the congenital absence of ganglion cells in some segment(s) of the colon. The muscle cannot make coordinated movements to propel stool in that section, most commonly leading to obstruction. The diagnosis and treatment for this disease require a clear identification of different region(s) of the myenteric plexus, where ganglion cells should be present, on the microscopic view of the tissue slide. While deep learning approaches, such as Convolutional Neural Networks, have performed very well in this task, they are often treated as black boxes, with minimal understanding gained from them, and may not conform to how a physician makes decisions. In this study, we propose a novel framework that integrates expert-derived textual concepts into a Contrastive Language-Image Pre-training-based vision-language model to guide plexus classification. Using prompts derived from expert sources (e.g., medical textbooks and papers) generated by large language models and reviewed by our team before being encoded with QuiltNet, our approach aligns clinically relevant semantic cues with visual features. Experimental results show that the proposed model demonstrated superior discriminative capability across different classification metrics as it outperformed CNN-based models, including VGG-19, ResNet-18, and ResNet-50; achieving an accuracy of 83.9%, a precision of 86.6%, and a specificity of 87.6%. These findings highlight the potential of multi-modal learning in histopathology and underscore the value of incorporating expert knowledge for more clinically relevant model outputs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。