GPT-5在眼科问答中表现最优,验证了推理强度对准确率的关键作用。
Performance of GPT-5 Frontier Models in Ophthalmology Question Answering
- 测试12种GPT-5配置,结合不同推理强度评估性能
- GPT-5-high准确率达96.5%,显著优于多数基线模型
- 提出可扩展的自动化评分框架,适用于眼科答案评估
大型语言模型(LLMs)如GPT-5具备先进推理能力,可能提升复杂医学问答任务表现。本文评估了OpenAI GPT-5系列12种配置(三种模型层级×四种推理强度设置),并与o1-high、o3-high及GPT-4o对比,使用260道来自美国眼科学会基础临床科学课程(BCSC)的闭源选择题。主要终点为选择题准确率;次要终点包括基于布拉德利-特里模型的头对头排名、使用参考锚定的成对大模型判官框架评估推理质量,以及基于令牌成本估算的准确率-成本权衡分析。GPT-5-high准确率达0.965(95%置信区间:0.942–0.985),显著优于所有GPT-5-nano变体(P < .001)、o1-high(P = .04)和GPT-4o(P < .001),但与o3-high(0.958,95% CI:0.931–0.981)无显著差异。GPT-5-high在准确率(比o3-high强1.66倍)和推理质量(强1.11倍)上均排名第一。成本-准确率分析识别出多个位于帕累托前沿的GPT-5配置,其中GPT-5-mini-low在低代价高表现间取得最佳平衡。本研究在高质量眼科数据集上基准化了GPT-5,揭示了推理努力对准确率的影响,并引入一种可扩展的自动评分框架,用于大模型生成答案与参考标准的规模化评估。
原文摘要 · Abstract (English)
Large language models (LLMs) such as GPT-5 integrate advanced reasoning capabilities that may improve performance on complex medical question-answering tasks. For this latest generation of reasoning models, the configurations that maximize both accuracy and cost-efficiency have yet to be established. We evaluated 12 configurations of OpenAI's GPT-5 series (three model tiers across four reasoning effort settings) alongside o1-high, o3-high, and GPT-4o, using 260 closed-access multiple-choice questions from the American Academy of Ophthalmology Basic Clinical Science Course (BCSC) dataset. The primary outcome was multiple-choice accuracy; secondary outcomes included head-to-head ranking via a Bradley-Terry model, rationale quality assessment using a reference-anchored, pairwise LLM-as-a-judge framework, and analysis of accuracy-cost trade-offs using token-based cost estimates. GPT-5-high achieved the highest accuracy (0.965; 95% CI, 0.942-0.985), outperforming all GPT-5-nano variants (P < .001), o1-high (P = .04), and GPT-4o (P < .001), but not o3-high (0.958; 95% CI, 0.931-0.981). GPT-5-high ranked first in both accuracy (1.66x stronger than o3-high) and rationale quality (1.11x stronger than o3-high). Cost-accuracy analysis identified several GPT-5 configurations on the Pareto frontier, with GPT-5-mini-low offering the most favorable low-cost, high-performance balance. These results benchmark GPT-5 on a high-quality ophthalmology dataset, demonstrate the influence of reasoning effort on accuracy, and introduce an autograder framework for scalable evaluation of LLM-generated answers against reference standards in ophthalmology.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。