用大模型集体评分系统替代医生评估医疗诊断,效果可靠且更高效。
Can LLMs Accurately Score Medical Diagnoses and Clinical Reasoning?

- 构建三个顶级大模型组成的评审团,对300例发展中国家病例进行评分。
- 校准后的大模型评分与专家评分高度一致,严重误诊风险更低。
- 可自动识别高风险诊断,帮助专家聚焦关键问题,提升评估效率。
使用专家医生小组评估医疗AI系统成本高、耗时长,促使研究者探索大语言模型(LLMs)作为替代评判工具。本文评估了一个由三个前沿AI模型组成的LLM评审团,在300个真实世界低收入和中等收入国家(LMIC)医院病例上对3334个诊断的评分表现。所有诊断(包括LLM生成和医生生成)均依据四个维度:诊断、鉴别诊断、临床推理及负向治疗风险,与专家小组诊断结果对比。通过与专家及独立重评小组的评分比较,分析误差指标、评分者间一致性、严重风险错误率,以及使用等熵回归进行事后校准的效果。结果显示:(i)未经校准的LLM评审团评分与专家评分保持有序一致性,但系统性偏低;(ii)LLM评审团的严重风险错误概率低于人类重评小组;(iii)结合LLM诊断的评审团可有效识别高风险诊断,实现针对性专家复审,提升评审效率;(iv)校准后的LLM评审团评分与主专家小组评分及排序高度一致;(v)LLM评审团无自偏好偏差,不会因诊断来自同一厂商或自身模型而给予更高或更低评分。上述结果表明,经校准的LLM评审团可作为医疗AI评测中专家评价的可信代理。未来需在其他临床场景中进一步验证。
原文摘要 · Abstract (English)
Evaluating medical AI systems using expert clinician panels is costly and slow, motivating the use of large language models (LLMs) as alternative adjudicators. Here, we evaluate an LLM Jury, composed of three frontier AI models, for scoring 3334 diagnoses on 300 real-world low- and middle-income country (LMIC) hospital cases. Both LLM- and clinician-generated diagnoses are scored against expert panel diagnoses across four dimensions: diagnosis, differential diagnosis, clinical reasoning, and negative treatment risk. The LLM Jury scores are compared with expert and independent re-scoring panel scores to assess error metrics, inter-rater agreement, severe-risk errors, and the effect of post hoc calibration using isotonic regression. In our data, we find that: (i) the uncalibrated LLM Jury scores preserve ordinal agreement with the expert clinician panel scores, but are systematically lower; (ii) the probability of severe-risk errors is lower for the LLM Jury than the human expert re-score panels; (iii) the LLM Jury combined with LLM diagnoses can be used to identify diagnoses at high risk of error, enabling targeted expert review and improved panel efficiency; (iv) the calibrated LLM Jury scores and rankings of diagnosing agents show excellent agreement with those of the primary expert panels; (v) LLM Jury models show no self-preference bias, they did not score diagnoses generated by their own underlying model or models from the same vendor more (or less) favourably than those generated by other models. Together, these results provide evidence that a calibrated LLM Jury is a trustworthy and reliable proxy for expert clinician evaluation in medical AI benchmarking. Confirming these findings in other clinical settings is an important direction for future work.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。