arXiv:2411.05897cs.CLcs.AI2024-11被引 8

LLMs在医学计算器推荐上不如医生,理解与知识仍有明显短板。

Humans and Large Language Models in Clinical Decision Support: A Study with Medical Calculators

  • 对比9个LLM与人类在35种医学计算器上的表现
  • 顶尖LLM准确率66.0%,人类达79.5%(100题子集)
  • 适合临床决策支持研究者参考,尤其关注人机协作

尽管大语言模型(LLMs)已在执业医师考试中评估过通用医学知识,但其在临床决策支持中的能力,如选择医学计算器,仍不明确。本研究评估了9个大语言模型,包括开源、专有和领域特定模型,使用35种临床计算器的1,009个多项选择题-答案对,并在100个问题子集上与人类进行对比。最高性能的LLM OpenAI o1 在该子集上准确率为66.0%(95%置信区间:56.7–75.3%),而两名人类标注员平均准确率为79.5%(95%置信区间:73.5–85.0%)。进一步评估显示,医学训练生与LLMs在风险分层和诊断等临床场景中推荐计算器的表现。错误分析表明,顶级LLM仍存在理解偏差(49.3%错误)和计算器知识不足(7.1%错误)。结果表明,目前LLMs在计算器推荐任务中不及人类。

原文摘要 · Abstract (English)

Although large language models (LLMs) have been assessed for general medical knowledge using licensing exams, their ability to support clinical decision-making, such as selecting medical calculators, remains uncertain. We assessed nine LLMs, including open-source, proprietary, and domain-specific models, with 1,009 multiple-choice question-answer pairs across 35 clinical calculators and compared LLMs to humans on a subset of questions. While the highest-performing LLM, OpenAI o1, provided an answer accuracy of 66.0% (CI: 56.7-75.3%) on the subset of 100 questions, two human annotators nominally outperformed LLMs with an average answer accuracy of 79.5% (CI: 73.5-85.0%). Ultimately, we evaluated medical trainees and LLMs in recommending medical calculators across clinical scenarios like risk stratification and diagnosis. With error analysis showing that the highest-performing LLMs continue to make mistakes in comprehension (49.3% of errors) and calculator knowledge (7.1% of errors), our findings highlight that LLMs are not superior to humans in calculator recommendation.

临床决策大模型评测医学应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。