arXiv:2503.12927cs.CVcs.AI2025-03被引 6

用图像+自动生成描述提升神经母细胞瘤分型准确率

MMLNB: Multi-Modal Learning for Neuroblastoma Subtyping Classification Assisted with Textual Description Generation

  • 图像与自动生成的文本描述联合分析
  • 比单模态模型准确率更高,验证了多模态融合有效性
  • 适合医学影像分析与可解释AI研究者参考

神经母细胞瘤(NB)是儿童癌症致死主因,其组织病理学表现差异大,需精准分型以指导预后和治疗。传统诊断依赖主观评估,耗时且不一致。为此,我们提出MMLNB,一种融合病理图像与生成文本描述的多模态学习模型。该方法分两阶段:首先微调视觉语言模型(VLM),提升病理感知文本生成能力;随后采用双分支结构分别提取视觉与文本特征,通过渐进式鲁棒多模态融合(PRMF)块实现特征融合,保障训练稳定性。实验表明,MMLNB优于单模态模型。消融实验证明多模态融合、微调及PRMF机制均至关重要。本研究构建了一个可扩展的AI数字病理框架,提升了NB分型的可靠性与可解释性。源代码见https://github.com/HovChen/MMLNB。

原文摘要 · Abstract (English)

Neuroblastoma (NB), a leading cause of childhood cancer mortality, exhibits significant histopathological variability, necessitating precise subtyping for accurate prognosis and treatment. Traditional diagnostic methods rely on subjective evaluations that are time-consuming and inconsistent. To address these challenges, we introduce MMLNB, a multi-modal learning (MML) model that integrates pathological images with generated textual descriptions to improve classification accuracy and interpretability. The approach follows a two-stage process. First, we fine-tune a Vision-Language Model (VLM) to enhance pathology-aware text generation. Second, the fine-tuned VLM generates textual descriptions, using a dual-branch architecture to independently extract visual and textual features. These features are fused via Progressive Robust Multi-Modal Fusion (PRMF) Block for stable training. Experimental results show that the MMLNB model is more accurate than the single modal model. Ablation studies demonstrate the importance of multi-modal fusion, fine-tuning, and the PRMF mechanism. This research creates a scalable AI-driven framework for digital pathology, enhancing reliability and interpretability in NB subtyping classification. Our source code is available at https://github.com/HovChen/MMLNB.

多模态病理分析AI医疗生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。