对比大模型与传统深度学习在脑出血分型中的表现,发现后者更准但前者更易解释。
Zero-Shot Multi-modal Large Language Model v.s. Supervised Deep Learning: A Comparative Study on CT-Based Intracranial Hemorrhage Subtyping
- 用提示工程引导多模态大模型处理CT图像的出血检测与分型任务
- 传统模型在二分类和分型上均优于大模型,大模型最高F1仅0.31
- 大模型虽不准但可生成语言解释,适合需要可解释性的临床场景
及时识别非增强CT上的颅内出血(ICH)亚型对预后判断和治疗决策至关重要,但因对比度低、边界模糊而具挑战性。本研究评估了零样本多模态大语言模型(MLLMs)与传统深度学习方法在ICH二分类及亚型分类中的表现。使用RSNA提供的192例NCCT数据集,比较GPT-4o、Gemini 2.0 Flash、Claude 3.5 Sonnet V2等MLLMs与ResNet50、Vision Transformer等传统模型。通过精心设计的提示引导模型完成出血存在判断、亚型分类、定位及体积估计任务。结果显示,在二分类任务中,传统深度学习模型全面优于MLLMs;在亚型分类中,MLLMs表现仍较差,其中Gemini 2.0 Flash的宏平均精确率为0.41,宏平均F1为0.31。结论表明,尽管MLLMs在交互性方面表现优异,其在ICH亚型分类的整体准确性仍低于深度神经网络,但其通过语言交互提升可解释性,显示出在医学影像分析中的潜力。未来工作将聚焦于模型优化与开发更精准的三维医学图像处理大模型。
原文摘要 · Abstract (English)
Introduction: Timely identification of intracranial hemorrhage (ICH) subtypes on non-contrast computed tomography is critical for prognosis prediction and therapeutic decision-making, yet remains challenging due to low contrast and blurring boundaries. This study evaluates the performance of zero-shot multi-modal large language models (MLLMs) compared to traditional deep learning methods in ICH binary classification and subtyping. Methods: We utilized a dataset provided by RSNA, comprising 192 NCCT volumes. The study compares various MLLMs, including GPT-4o, Gemini 2.0 Flash, and Claude 3.5 Sonnet V2, with conventional deep learning models, including ResNet50 and Vision Transformer. Carefully crafted prompts were used to guide MLLMs in tasks such as ICH presence, subtype classification, localization, and volume estimation. Results: The results indicate that in the ICH binary classification task, traditional deep learning models outperform MLLMs comprehensively. For subtype classification, MLLMs also exhibit inferior performance compared to traditional deep learning models, with Gemini 2.0 Flash achieving an macro-averaged precision of 0.41 and a macro-averaged F1 score of 0.31. Conclusion: While MLLMs excel in interactive capabilities, their overall accuracy in ICH subtyping is inferior to deep networks. However, MLLMs enhance interpretability through language interactions, indicating potential in medical imaging analysis. Future efforts will focus on model refinement and developing more precise MLLMs to improve performance in three-dimensional medical image processing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。