GPT-5在医学影像推理任务中表现远超GPT-4o,尤其在胸部和肺部等复杂区域。
Benchmarking GPT-5 for Zero-Shot Multimodal Medical Reasoning in Radiology and Radiation Oncology
- 零样本评估GPT-5在多模态医学任务中的表现
- 在胸部区域准确率比GPT-4o高20.00%,物理考题达90.7%正确率
- 适合医疗影像与放疗物理领域的专家辅助决策
放射科、放疗学和医学物理领域需在高风险条件下整合医学图像、文本报告与定量数据进行决策。随着GPT-5的推出,亟需评估大模型进展是否带来实际提升。我们针对GPT-5及其小型变体(GPT-5-mini、GPT-5-nano)与GPT-4o,在三个代表性任务上开展零样本评估:(1) VQA-RAD,放射影像视觉问答基准;(2) SLAKE,测试跨模态对齐的多语言语义标注VQA数据集;(3) 150道医学物理执照考试风格的多选题,涵盖治疗计划、剂量学、成像与质量保障。在所有数据集上,GPT-5均达到最高准确率,相较GPT-4o在胸纵隔区域提升20.00%,肺部问题提升13.60%,脑组织解析提升11.44%。在物理考题中,GPT-5准确率达90.7%(136/150),超过预估人类及格线,而GPT-4o为78.0%。结果表明,GPT-5在图像引导推理与领域特定数值求解中均实现稳定且显著优于GPT-4o的表现,展现出其在医学影像与治疗物理专家工作流中的增强潜力。
原文摘要 · Abstract (English)
Radiology, radiation oncology, and medical physics require decision-making that integrates medical images, textual reports, and quantitative data under high-stakes conditions. With the introduction of GPT-5, it is critical to assess whether recent advances in large multimodal models translate into measurable gains in these safety-critical domains. We present a targeted zero-shot evaluation of GPT-5 and its smaller variants (GPT-5-mini, GPT-5-nano) against GPT-4o across three representative tasks. We present a targeted zero-shot evaluation of GPT-5 and its smaller variants (GPT-5-mini, GPT-5-nano) against GPT-4o across three representative tasks: (1) VQA-RAD, a benchmark for visual question answering in radiology; (2) SLAKE, a semantically annotated, multilingual VQA dataset testing cross-modal grounding; and (3) a curated Medical Physics Board Examination-style dataset of 150 multiple-choice questions spanning treatment planning, dosimetry, imaging, and quality assurance. Across all datasets, GPT-5 achieved the highest accuracy, with substantial gains over GPT-4o up to +20.00% in challenging anatomical regions such as the chest-mediastinal, +13.60% in lung-focused questions, and +11.44% in brain-tissue interpretation. On the board-style physics questions, GPT-5 attained 90.7% accuracy (136/150), exceeding the estimated human passing threshold, while GPT-4o trailed at 78.0%. These results demonstrate that GPT-5 delivers consistent and often pronounced performance improvements over GPT-4o in both image-grounded reasoning and domain-specific numerical problem-solving, highlighting its potential to augment expert workflows in medical imaging and therapeutic physics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。