arXiv:2506.02987cs.CLcs.AI2025-06被引 2

测试4大顶级大模型在全科医生考试中的表现,均远超人类医生平均水平。

Performance of leading large language models in May 2025 in Membership of the Royal College of General Practitioners-style examination questions: a cross-sectional analysis

  • 用真实全科医生考试题测试o3、Claude Opus 4等4个模型
  • o3得分99.0%,其他模型95.0%,远超人类平均73.0%
  • 适合关注AI辅助医疗、医学教育的从业者参考

背景:大型语言模型(LLMs)在临床实践中展现出巨大潜力。然而,除ChatGPT4及其前代模型外,很少有领先的大模型,尤其是具备强推理能力的模型,被用于医学专科考试,特别是在全科医学领域。本文旨在评估截至2025年5月的领先大模型(o3、Claude Opus 4、Grok3和Gemini 2.5 Pro)在全科医学教育中的表现,具体为回答英国皇家全科医师学会(MRCGP)风格的试题。方法:在2025年5月25日,使用o3、Claude Opus 4、Grok3和Gemini 2.5 Pro分别回答100道随机选取的英国皇家全科医师学会GP SelfTest中的多选题,题目包含文本信息、实验室结果及临床图像。每个模型以英国全科医生身份作答,并获得完整题干信息,每题仅尝试一次。答案由GP SelfTest提供的正确答案评分。结果:o3、Claude Opus 4、Grok3和Gemini 2.5 Pro的总分为99.0%、95.0%、95.0%和95.0%。相同题目的平均人类得分仅为73.0%。讨论:所有模型表现优异,显著超过一般全科医生及住院医师水平。o3表现最佳,其余模型表现相近且未明显低于o3。这些发现强化了大模型,特别是针对全科临床数据训练的推理模型,在支持初级保健服务中的应用前景。

原文摘要 · Abstract (English)

Background: Large language models (LLMs) have demonstrated substantial potential to support clinical practice. Other than Chat GPT4 and its predecessors, few LLMs, especially those of the leading and more powerful reasoning model class, have been subjected to medical specialty examination questions, including in the domain of primary care. This paper aimed to test the capabilities of leading LLMs as of May 2025 (o3, Claude Opus 4, Grok3, and Gemini 2.5 Pro) in primary care education, specifically in answering Member of the Royal College of General Practitioners (MRCGP) style examination questions. Methods: o3, Claude Opus 4, Grok3, and Gemini 2.5 Pro were tasked to answer 100 randomly chosen multiple choice questions from the Royal College of General Practitioners GP SelfTest on 25 May 2025. Questions included textual information, laboratory results, and clinical images. Each model was prompted to answer as a GP in the UK and was provided with full question information. Each question was attempted once by each model. Responses were scored against correct answers provided by GP SelfTest. Results: The total score of o3, Claude Opus 4, Grok3, and Gemini 2.5 Pro was 99.0%, 95.0%, 95.0%, and 95.0%, respectively. The average peer score for the same questions was 73.0%. Discussion: All models performed remarkably well, and all substantially exceeded the average performance of GPs and GP registrars who had answered the same questions. o3 demonstrated the best performance, while the performances of the other leading models were comparable with each other and were not substantially lower than that of o3. These findings strengthen the case for LLMs, particularly reasoning models, to support the delivery of primary care, especially those that have been specifically trained on primary care clinical data.

大模型医疗AI全科医学评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。