GPT-5在医学多模态推理中超越人类专家,实现精准诊断决策。
Capabilities of GPT-5 on Multimodal Medical Reasoning
- 采用零样本思维链框架,统一评估文本与视觉问答任务。
- 在多模态医学问答上比GPT-4o提升29.26%推理能力,超人类专家24.23%。
- 适合医疗决策支持系统研发者及临床人工智能研究者参考。
大型语言模型的进展使通用系统能在无需大量微调的情况下完成复杂领域推理。在医学领域,决策常需整合患者叙述、结构化数据和医学影像等异构信息。本研究将GPT-5定位为通用多模态医学推理器,采用统一协议系统评估其零样本思维链推理性能,涵盖文本问答与视觉问答任务。我们在MedQA、MedXpertQA(文本与多模态)、MMLU医学子集、USMLE自测题及VQA-RAD等标准数据集上对比GPT-5、GPT-5-mini、GPT-5-nano与GPT-4o-2024-11-20。结果表明,GPT-5在所有基准上均优于基线,达到当前最优准确率,在多模态推理中表现显著。在MedXpertQA MM上,其推理与理解得分分别较GPT-4o提升+29.26%和+26.18%,且在推理与理解上分别超过未持证人类专家+24.23%和+29.40%。而GPT-4o多数维度仍低于人类专家。典型案例显示,GPT-5能融合视觉与文本线索形成连贯诊断链,推荐高风险干预措施。结果显示,GPT-5在控制性多模态推理基准上已从人类相当跃升至超越人类专家水平,对下一代临床决策支持系统设计具有重要启示。
原文摘要 · Abstract (English)
Recent advances in large language models (LLMs) have enabled general-purpose systems to perform increasingly complex domain-specific reasoning without extensive fine-tuning. In the medical domain, decision-making often requires integrating heterogeneous information sources, including patient narratives, structured data, and medical images. This study positions GPT-5 as a generalist multimodal reasoner for medical decision support and systematically evaluates its zero-shot chain-of-thought reasoning performance on both text-based question answering and visual question answering tasks under a unified protocol. We benchmark GPT-5, GPT-5-mini, GPT-5-nano, and GPT-4o-2024-11-20 against standardized splits of MedQA, MedXpertQA (text and multimodal), MMLU medical subsets, USMLE self-assessment exams, and VQA-RAD. Results show that GPT-5 consistently outperforms all baselines, achieving state-of-the-art accuracy across all QA benchmarks and delivering substantial gains in multimodal reasoning. On MedXpertQA MM, GPT-5 improves reasoning and understanding scores by +29.26% and +26.18% over GPT-4o, respectively, and surpasses pre-licensed human experts by +24.23% in reasoning and +29.40% in understanding. In contrast, GPT-4o remains below human expert performance in most dimensions. A representative case study demonstrates GPT-5's ability to integrate visual and textual cues into a coherent diagnostic reasoning chain, recommending appropriate high-stakes interventions. Our results show that, on these controlled multimodal reasoning benchmarks, GPT-5 moves from human-comparable to above human-expert performance. This improvement may substantially inform the design of future clinical decision-support systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。