arXiv:2412.01306cs.AIcs.CV2024-12被引 5

用LLaMA II模型融合影像与文本,提升疾病分类准确率

Multimodal Medical Disease Classification with LLaMA II

  • 以LLaMA II为骨干,测试不同模态融合架构
  • 早期融合使平均AUC达97.10%,优于晚期融合的96.67%
  • 可迁移至其他医疗多模态数据,适合医学AI研究者

医疗患者数据具有多模态特性,如影像、文本、年龄、性别及病理数据等。利用深度学习处理并整合此类多模态信息,在诊断与治疗规划中具有巨大潜力。本文基于开放数据集OpenI(包含2D胸部X光片与临床报告)重新训练了一种基于Transformer的多模态模型,重点研究文本与视觉信息的融合策略。采用以LLaMA II为骨干的多种架构进行对比。结果显示,模态特异性特征的早期融合表现更优,最佳模型达到97.10%的平均AUC,高于深层晚期融合的最佳结果(96.67%)。两种方法均优于此前在相同数据集上测试的分类模型。所提出的多模态架构可轻松适配其他多模态数据集,便于后续研究,尤其适用于医学人工智能领域。

原文摘要 · Abstract (English)

Medical patient data is always multimodal. Images, text, age, gender, histopathological data are only few examples for different modalities in this context. Processing and integrating this multimodal data with deep learning based methods is of utmost interest due to its huge potential for medical procedure such as diagnosis and patient treatment planning. In this work we retrain a multimodal transformer-based model for disease classification. To this end we use the text-image pair dataset from OpenI consisting of 2D chest X-rays associated with clinical reports. Our focus is on fusion methods for merging text and vision information extracted from medical datasets. Different architecture structures with a LLaMA II backbone model are tested. Early fusion of modality specific features creates better results with the best model reaching 97.10% mean AUC than late fusion from a deeper level of the architecture (best model: 96.67% mean AUC). Both outperform former classification models tested on the same multimodal dataset. The newly introduced multimodal architecture can be applied to other multimodal datasets with little effort and can be easily adapted for further research, especially, but not limited to, the field of medical AI.

多模态医学图像LLaMA疾病分类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。