arXiv:2501.13949cs.CLcs.AI2025-01被引 3

OpenAI o1在眼科领域表现优异,但推理能力仍有提升空间。

Can OpenAI o1 Reason Well in Ophthalmology? A 6,990-Question Head-to-Head Evaluation Study

  • 用6990道眼科题对比o1与其他模型,评估其表现与推理能力。
  • o1准确率88%,在白内障和青光眼领域排名第一,整体排名第二。
  • 长答案题目中表现更优,适合需要深度推理的医学场景。

本研究通过6,990道来自MedMCQA的眼科问题,评估了OpenAI o1与其他五个大语言模型的表现。结果显示,o1在准确率(0.88)和宏平均F1分数上最高,但在基于文本生成的推理能力评测中排名第三。在子领域中,o1在“晶状体”和“青光眼”方面排名第一,但在“角膜及外眼疾病”“玻璃体与视网膜疾病”以及“眼整形与眶部疾病”三个领域落后于GPT-4o。亚组分析表明,o1在需要较长真实答案的问题中表现更优。结果表明,o1的推理增强机制尚未完全适用于眼科专业领域,提示需进行领域特异性优化以提升专业场景下的性能。

原文摘要 · Abstract (English)

Question: What is the performance and reasoning ability of OpenAI o1 compared to other large language models in addressing ophthalmology-specific questions? Findings: This study evaluated OpenAI o1 and five LLMs using 6,990 ophthalmological questions from MedMCQA. O1 achieved the highest accuracy (0.88) and macro-F1 score but ranked third in reasoning capabilities based on text-generation metrics. Across subtopics, o1 ranked first in ``Lens'' and ``Glaucoma'' but second to GPT-4o in ``Corneal and External Diseases'', ``Vitreous and Retina'' and ``Oculoplastic and Orbital Diseases''. Subgroup analyses showed o1 performed better on queries with longer ground truth explanations. Meaning: O1's reasoning enhancements may not fully extend to ophthalmology, underscoring the need for domain-specific refinements to optimize performance in specialized fields like ophthalmology.

眼科大模型推理能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。