arXiv:2505.23503eess.IVcs.AI2025-05被引 8

比较LLM与CNN在医学影像诊断中的表现,发现优化后LLM可显著提升效果。

Can Large Language Models Challenge CNNs in Medical Image Analysis?

  • 用图像+文本的多模态框架对比CNN与LLM的诊断能力
  • 优化后的LLM在准确率和效率上接近甚至超越CNN
  • 适合关注AI医疗系统选型与能效的临床研究者

本研究提出一种多模态AI框架,用于精准分类医学诊断图像。基于公开数据集,该系统对比了卷积神经网络(CNN)与不同大型语言模型(LLMs)的性能。深入的比较分析揭示了诊断准确性、执行效率及环境影响的关键差异。评估指标包括准确率、F1分数、平均执行时间、平均能耗及估算的$CO_2$排放量。结果表明,尽管基于CNN的模型在多数情况下优于融合图像与上下文信息的多模态方法,但在对LLM施加额外过滤后,性能可获得显著提升。这些发现突显了多模态AI系统在提升临床诊断可靠性、效率与可扩展性方面的变革潜力。

原文摘要 · Abstract (English)

This study presents a multimodal AI framework designed for precisely classifying medical diagnostic images. Utilizing publicly available datasets, the proposed system compares the strengths of convolutional neural networks (CNNs) and different large language models (LLMs). This in-depth comparative analysis highlights key differences in diagnostic performance, execution efficiency, and environmental impacts. Model evaluation was based on accuracy, F1-score, average execution time, average energy consumption, and estimated $CO_2$ emission. The findings indicate that although CNN-based models can outperform various multimodal techniques that incorporate both images and contextual information, applying additional filtering on top of LLMs can lead to substantial performance gains. These findings highlight the transformative potential of multimodal AI systems to enhance the reliability, efficiency, and scalability of medical diagnostics in clinical settings.

医学影像多模态大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。