用临床专家设计的评分标准,发现大模型在关键医疗决策上表现仍差。
A rubric-based controlled comparison of frontier language models on expert-authored clinical reasoning tasks
- 构建5个专科临床场景,配以专家制定的加权评分细则。
- 高权重关键题仅32%-42%被正确回答,低权重题达80%-90%。
- 结果揭示模型优先级错位,适合评估医疗AI可靠性。
多选医学评测已趋于饱和,近期基于评分标准的评估(如HealthBench)表明开放式临床表现仍未解决——其‘难’子集最高得分仍为32%。本文构建一个小型、刻意困难的评估数据集,包含五个由临床医生撰写的临床情境,覆盖麻醉学、内科/家庭医学、急诊医学和妇产科四个专科,每个任务均配有原子化、加权、互斥完备(MECE)的评分标准(每任务25-62项,共184项),依据临床医生撰写的黄金答案制定。我们评估了三款前沿模型:GPT 5.4、Claude Opus 4.7和Gemini 3.1 Pro。平均评分通过率分别为0.47(Claude)、0.39(GPT)和0.37(Gemini)。核心发现是临床优先级反转:最高权重(权重5,关键)标准仅通过32.4%-41.7%,而低权重(权重1)标准通过率达80%-90%。108项关键(权重5)标准中,56项(52%)无一模型满足。三位LLM自评器对552项标准的专家判定标签复现率达92.8%-94.7%。本研究定位为方法与初步发现贡献:五个任务展示了一条可扩展、有依据的开发路径,可用于构建大规模基准。
原文摘要 · Abstract (English)
Multiple-choice medical benchmarks are increasingly saturated, and recent rubric-based evaluations such as HealthBench have shown that open-ended clinical performance is far from solved - its "Hard" subset top score remains 32%. We present a small, deliberately difficult evaluation dataset of five clinician-authored clinical scenarios spanning four specialties (anaesthesia, internal/family medicine, emergency medicine, and obstetrics), each accompanied by an atomic, weighted, MECE rubric (25-62 criteria per task; 184 criteria total) authored from a clinician-drafted golden answer. We evaluate three frontier models: GPT 5.4, Claude Opus 4.7, and Gemini 3.1 Pro. Mean rubric pass rates were 0.47 (Claude), 0.39 (GPT), and 0.37 (Gemini). The central finding is an inversion of clinical priority: the highest-weighted (weight-5, critical) criteria passed at only 32.4-41.7%, while low-stakes weight-1 criteria passed at 80-90%. 56 of 108 critical (weight-5) criteria (52%) were satisfied by no model. Three LLM autoraters reproduced expert met/not-met labels on 92.8-94.7% of 552 graded criteria. We position this as a methods-and-preliminary-findings contribution: the five tasks demonstrate a scalable, defensible pipeline ready to develop into a large-scale benchmark.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。