评测4款新推理模型在眼科考试题上的表现,发现o1和DeepSeek-R1最准。
Benchmarking Next-Generation Reasoning-Focused Large Language Models in Ophthalmology: A Head-to-Head Evaluation on 5,888 Items
- 用5888道眼科题零样本测试4个推理型大模型
- o1和DeepSeek-R1准确率最高,达90.2%和88.8%
- Gemini 2.0最快,DeepSeek-R1最慢但推理最完整
近期聚焦推理的大语言模型(LLMs)从通用模型转向复杂决策能力,这对医学领域至关重要。然而其在眼科等专业领域的表现仍不明确。本研究全面评估并比较了四种新开发的推理型大模型:DeepSeek-R1、OpenAI o1、o3-mini 和 Gemini 2.0 Flash-Thinking。所有模型在零样本设置下使用来自 MedMCQA 数据集的 5,888 道眼科多选题进行测试。定量评估包括准确率、Macro-F1 及五项文本生成指标(ROUGE-L、METEOR、BERTScore、BARTScore、AlignScore),均与真实推理答案对比。对随机选取的 100 道题记录平均推理时间。此外,两名持证眼科医生对差分诊断类问题的回答进行定性评估,考察清晰度、完整性与推理结构。结果显示,o1(0.902)和 DeepSeek-R1(0.888)准确率最高,o1 在 Macro-F1 上也领先(0.900)。各模型在文本生成指标上表现各异:o3-mini 在 ROUGE-L 上最优(0.151),o1 在 METEOR(0.232)领先,DeepSeek-R1 与 o3-mini 并列 BERTScore(0.673),DeepSeek-R1(-4.105)与 Gemini 2.0 Flash-Thinking(-4.127)在 BARTScore 表现最佳,o3-mini(0.181)和 o1(0.176)在 AlignScore 上领先。推理时间差异显著:DeepSeek-R1 最慢(40.4 秒),Gemini 2.0 Flash-Thinking 最快(6.7 秒)。定性分析显示,DeepSeek-R1 与 Gemini 2.0 Flash-Thinking 提供更详细、全面的中间推理过程,而 o1 与 o3-mini 的回答更简洁、总结性强。
原文摘要 · Abstract (English)
Recent advances in reasoning-focused large language models (LLMs) mark a shift from general LLMs toward models designed for complex decision-making, a crucial aspect in medicine. However, their performance in specialized domains like ophthalmology remains underexplored. This study comprehensively evaluated and compared the accuracy and reasoning capabilities of four newly developed reasoning-focused LLMs, namely DeepSeek-R1, OpenAI o1, o3-mini, and Gemini 2.0 Flash-Thinking. Each model was assessed using 5,888 multiple-choice ophthalmology exam questions from the MedMCQA dataset in zero-shot setting. Quantitative evaluation included accuracy, Macro-F1, and five text-generation metrics (ROUGE-L, METEOR, BERTScore, BARTScore, and AlignScore), computed against ground-truth reasonings. Average inference time was recorded for a subset of 100 randomly selected questions. Additionally, two board-certified ophthalmologists qualitatively assessed clarity, completeness, and reasoning structure of responses to differential diagnosis questions.O1 (0.902) and DeepSeek-R1 (0.888) achieved the highest accuracy, with o1 also leading in Macro-F1 (0.900). The performance of models across the text-generation metrics varied: O3-mini excelled in ROUGE-L (0.151), o1 in METEOR (0.232), DeepSeek-R1 and o3-mini tied for BERTScore (0.673), DeepSeek-R1 (-4.105) and Gemini 2.0 Flash-Thinking (-4.127) performed best in BARTScore, while o3-mini (0.181) and o1 (0.176) led AlignScore. Inference time across the models varied, with DeepSeek-R1 being slowest (40.4 seconds) and Gemini 2.0 Flash-Thinking fastest (6.7 seconds). Qualitative evaluation revealed that DeepSeek-R1 and Gemini 2.0 Flash-Thinking tended to provide detailed and comprehensive intermediate reasoning, whereas o1 and o3-mini displayed concise and summarized justifications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。