测试前沿AI在复杂放射诊断中的表现,发现其远不如专业医生。
Radiology's Last Exam (RadLE): Benchmarking Frontier Multimodal AI Against Human Experts and a Taxonomy of Visual Reasoning Errors in Radiology
- 用50个高难度影像病例对比AI与放射科医生诊断能力。
- 最佳AI模型GPT-5准确率仅30%,远低于医生的83%。
- 提出视觉推理错误分类体系,帮助改进AI诊断可靠性。
通用多模态AI系统如大语言模型(LLMs)和视觉语言模型(VLMs)正被临床医生和患者通过消费级聊天机器人用于医学影像解读。多数宣称达到专家水平的评估基于包含常见病灶的公开数据集,对前沿模型在高难度诊断案例中的严格评估仍有限。我们构建了一个包含50个专家级“点诊断”病例的试点基准,涵盖多种影像模态,评估前沿AI模型在真实场景下的表现。测试了五种主流前沿AI模型(OpenAI o3、GPT-5、Gemini 2.5 Pro、Grok-4、Claude Opus 4.1)通过其原生网页接口的推理能力,由盲评专家打分,并进行三次独立复现。此外,对GPT-5在不同推理模式下的表现进行了评估。结果表明,执业放射科医生诊断准确率达83%,显著优于住院医师(45%)及所有AI模型(最高为GPT-5的30%)。可靠性方面,GPT-5和o3表现良好,Gemini 2.5 Pro和Grok-4中等,Claude Opus 4.1较差。研究揭示前沿模型在复杂诊断中仍远逊于人类专家,警示不应未经监督地应用于临床。同时,我们分析了推理过程并提出一套实用的视觉推理错误分类体系,以助理解模型失效机制,推动评估标准完善与更稳健模型开发。
原文摘要 · Abstract (English)
Generalist multimodal AI systems such as large language models (LLMs) and vision language models (VLMs) are increasingly accessed by clinicians and patients alike for medical image interpretation through widely available consumer-facing chatbots. Most evaluations claiming expert level performance are on public datasets containing common pathologies. Rigorous evaluation of frontier models on difficult diagnostic cases remains limited. We developed a pilot benchmark of 50 expert-level "spot diagnosis" cases across multiple imaging modalities to evaluate the performance of frontier AI models against board-certified radiologists and radiology trainees. To mirror real-world usage, the reasoning modes of five popular frontier AI models were tested through their native web interfaces, viz. OpenAI o3, OpenAI GPT-5, Gemini 2.5 Pro, Grok-4, and Claude Opus 4.1. Accuracy was scored by blinded experts, and reproducibility was assessed across three independent runs. GPT-5 was additionally evaluated across various reasoning modes. Reasoning quality errors were assessed and a taxonomy of visual reasoning errors was defined. Board-certified radiologists achieved the highest diagnostic accuracy (83%), outperforming trainees (45%) and all AI models (best performance shown by GPT-5: 30%). Reliability was substantial for GPT-5 and o3, moderate for Gemini 2.5 Pro and Grok-4, and poor for Claude Opus 4.1. These findings demonstrate that advanced frontier models fall far short of radiologists in challenging diagnostic cases. Our benchmark highlights the present limitations of generalist AI in medical imaging and cautions against unsupervised clinical use. We also provide a qualitative analysis of reasoning traces and propose a practical taxonomy of visual reasoning errors by AI models for better understanding their failure modes, informing evaluation standards and guiding more robust model development.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。