GPT-5在临床多模态推理中表现显著提升,尤其在影像辅助诊断上超越前代模型。
Evaluating GPT-5 as a Multimodal Clinical Reasoner: A Landscape Commentary
- 采用零样本链式思考框架,评估GPT-5系列在医学问答与影像理解中的综合推理能力。
- 在医学文本推理上提升超25个百分点,在乳腺钼靶细粒度分析中比GPT-4o高10%-40%。
- 虽接近临床认知模式,但在神经放射和乳腺影像中仍逊于专业模型,不宜替代专用系统。
从任务特定AI向通用基础模型的演进,引发对模型在临床医学中整合推理能力的根本质疑——诊断需融合模糊病史、实验室数据与多模态影像。本文首次开展控制性横断面评估,对比GPT-5家族(GPT-5、GPT-5 Mini、GPT-5 Nano)与前代GPT-4o在多样化临床任务中的表现,涵盖医学教育考试、文本推理基准及神经放射学、数字病理学、乳腺钼靶的视觉问答任务,均采用标准化零样本链式思考协议。GPT-5在专家级文本推理上实现显著提升,MedXpertQA绝对得分提升超25个百分点;在多模态融合中,有效利用增强推理能力将不确定病史与具体影像证据结合,在多数视觉问答任务中达到领先或可比水平,尤其在乳腺钼靶细粒度病变识别中较GPT-4o高出10%-40%。然而,神经放射学任务平均准确率仅为44%,且在乳腺钼靶领域落后于专用系统(其准确率超80%,而GPT-5为52%-64%)。结果表明,尽管GPT-5在迈向集成多模态临床推理方面迈出重要一步,模拟了医生以客观发现校准模糊信息的认知过程,但通用模型尚未能替代专精系统在高度专业化、感知敏感的任务中发挥作用。
原文摘要 · Abstract (English)
The transition from task-specific artificial intelligence toward general-purpose foundation models raises fundamental questions about their capacity to support the integrated reasoning required in clinical medicine, where diagnosis demands synthesis of ambiguous patient narratives, laboratory data, and multimodal imaging. This landscape commentary provides the first controlled, cross-sectional evaluation of the GPT-5 family (GPT-5, GPT-5 Mini, GPT-5 Nano) against its predecessor GPT-4o across a diverse spectrum of clinically grounded tasks, including medical education examinations, text-based reasoning benchmarks, and visual question-answering in neuroradiology, digital pathology, and mammography using a standardized zero-shot chain-of-thought protocol. GPT-5 demonstrated substantial gains in expert-level textual reasoning, with absolute improvements exceeding 25 percentage-points on MedXpertQA. When tasked with multimodal synthesis, GPT-5 effectively leveraged this enhanced reasoning capacity to ground uncertain clinical narratives in concrete imaging evidence, achieving state-of-the-art or competitive performance across most VQA benchmarks and outperforming GPT-4o by margins of 10-40% in mammography tasks requiring fine-grained lesion characterization. However, performance remained moderate in neuroradiology (44% macro-average accuracy) and lagged behind domain-specific models in mammography, where specialized systems exceed 80% accuracy compared to GPT-5's 52-64%. These findings indicate that while GPT-5 represents a meaningful advance toward integrated multimodal clinical reasoning, mirroring the clinician's cognitive process of biasing uncertain information with objective findings, generalist models are not yet substitutes for purpose-built systems in highly specialized, perception-critical tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。