22个大模型在西班牙医考上比拼,专精模型表现更优。
Evaluating Large Language Models on the Spanish Medical Intern Resident (MIR) Examination 2024/2025:A Comparative Analysis of Clinical Reasoning and Knowledge Application
- 对比22个大模型在西班牙医考上的临床推理与多模态能力
- 专精模型准确率显著高于通用模型,尤其在细粒度医学问题上
- 适合医疗AI研究者、医学教育者及临床决策系统开发者
本研究对22个大型语言模型(LLMs)在2024/2025年西班牙医学住院医师资格考试(MIR)上的表现进行了对比评估,重点关注临床推理、领域专业知识和多模态处理能力。MIR考试包含210道选择题,部分题目需图像解读,是检验事实记忆与复杂临床问题解决能力的严苛基准。评估涵盖GPT-4、Claude、LLaMA、Gemini等通用模型,以及基于西班牙医疗数据微调的专用模型Miri Pro,还包括近期入局的Deepseek和Grok,后者在视觉与语义分析任务中表现突出。结果显示,尽管通用模型整体表现稳健,但微调模型在应对细微领域挑战时始终更优;两轮考试间出现轻微成绩下降,归因于题目优化以减少对记忆依赖。研究凸显了领域微调与多模态融合在推动医疗AI应用中的变革潜力,并强调将大模型融入医学教育、培训与临床决策时,需兼顾自动化推理与伦理、情境感知判断。
原文摘要 · Abstract (English)
This study presents a comparative evaluation of 22 large language models LLMs on the Spanish Medical Intern Resident MIR examinations for 2024 and 2025 with a focus on clinical reasoning domain specific expertise and multimodal processing capabilities The MIR exam consisting of 210 multiple choice questions some requiring image interpretation serves as a stringent benchmark for assessing both factual recall and complex clinical problem solving skills Our investigation encompasses general purpose models such as GPT4 Claude LLaMA and Gemini as well as specialized fine tuned systems like Miri Pro which leverages proprietary Spanish healthcare data to excel in medical contexts Recent market entries Deepseek and Grok have further enriched the evaluation landscape particularly for tasks that demand advanced visual and semantic analysis The findings indicate that while general purpose LLMs perform robustly overall fine tuned models consistently achieve superior accuracy especially in addressing nuanced domain specific challenges A modest performance decline observed between the two exam cycles appears attributable to the implementation of modified questions designed to mitigate reliance on memorization The results underscore the transformative potential of domain specific fine tuning and multimodal integration in advancing medical AI applications They also highlight critical implications for the future integration of LLMs into medical education training and clinical decision making emphasizing the importance of balancing automated reasoning with ethical and context aware judgment
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。