arXiv:2509.18015cs.CVcs.AI2025-09被引 2

评测大模型在胸片病灶定位能力,发现性能远低于专业模型和医生。

Beyond Diagnosis: Evaluating Multimodal LLMs for Pathology Localization in Chest Radiographs

  • 用网格提示法让多模态大模型输出病灶坐标,评估其空间定位能力。
  • GPT-5定位准确率49.7%,仍低于专业模型的59.9%和医生的80.1%。
  • 模型对固定位置病灶表现较好,但对可变位置病灶定位不准,适合辅助而非替代。

近期研究显示前沿大语言模型及其多模态版本在医学问答和诊断任务中表现优异,具备广泛临床应用潜力。然而,医学图像解读的关键环节不仅是诊断,还包括病灶定位。定位能力评估兼具临床与教育意义,并能反映模型对解剖结构与疾病的空间理解。本文系统评估了两种通用多模态大模型(GPT-4、GPT-5)和一个领域专用模型(MedGemma)在胸片病灶定位上的表现,采用基于空间网格的提示策略,要求模型输出坐标。在CheXlocalize数据集上,九类病灶平均定位准确率分别为:GPT-5为49.7%,GPT-4为39.1%,MedGemma为17.7%,均显著低于任务专用CNN基线(59.9%)和放射科医生基准(80.1%)。尽管整体表现有限,误差分析显示,GPT-5预测大多位于解剖学合理区域,但位置不够精确;GPT-4对固定解剖位置病灶表现良好,却难以处理空间可变病灶,且常出现解剖学不合理预测;MedGemma整体表现最差,但在少量示例提示下有所提升。结果表明当前多模态大模型在医学影像中仍有局限,需与专用工具结合以实现可靠应用。

原文摘要 · Abstract (English)

Recent work has shown promising performance of frontier large language models (LLMs) and their multimodal counterparts in medical quizzes and diagnostic tasks, highlighting their potential for broad clinical utility given their accessible, general-purpose nature. However, beyond diagnosis, a fundamental aspect of medical image interpretation is the ability to localize pathological findings. Evaluating localization not only has clinical and educational relevance but also provides insight into a model's spatial understanding of anatomy and disease. Here, we systematically assess two general-purpose MLLMs (GPT-4 and GPT-5) and a domain-specific model (MedGemma) in their ability to localize pathologies on chest radiographs, using a prompting pipeline that overlays a spatial grid and elicits coordinate-based predictions. Averaged across nine pathologies in the CheXlocalize dataset, GPT-5 exhibited a localization accuracy of 49.7%, followed by GPT-4 (39.1%) and MedGemma (17.7%), all lower than a task-specific CNN baseline (59.9%) and a radiologist benchmark (80.1%). Despite modest performance, error analysis revealed that GPT-5's predictions were largely in anatomically plausible regions, just not always precisely localized. GPT-4 performed well on pathologies with fixed anatomical locations, but struggled with spatially variable findings and exhibited anatomically implausible predictions more frequently. MedGemma demonstrated the lowest performance on all pathologies, but showed improvements when provided examples through few shot prompting. Our findings highlight both the promise and limitations of current MLLMs in medical imaging and underscore the importance of integrating them with task-specific tools for reliable use.

病灶定位多模态模型医学影像大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。