arXiv:2503.17599cs.CLcs.AI2025-03被引 4

用真实临床标准评估大模型能否当全科医生

Evaluating Clinical Competencies of Large Language Models with a General Practice Benchmark

  • 构建专家标注的全科临床任务数据集
  • 10个主流大模型均需人类监督不可独立行医
  • 强调模型需针对全科职责专门优化

大语言模型在全科医学中展现出巨大潜力,但现有评估多依赖考试式问答,缺乏与真实临床职责对齐的能力框架。本文提出新评估体系,构建由领域专家按常规临床标准标注的全科医学基准(GPBench),评估十种先进大模型的实际能力。结果表明,当前大模型尚不适合在临床全科中自主部署,所有实际应用均需持续人工监督;针对全科医生日常职责的专门优化仍至关重要。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have demonstrated considerable potential in general practice. However, existing benchmarks and evaluation frameworks primarily depend on exam-style or simplified question-answer formats, lacking a competency-based structure aligned with the real-world clinical responsibilities encountered in general practice. Consequently, the extent to which LLMs can reliably fulfill the duties of general practitioners (GPs) remains uncertain. In this work, we propose a novel evaluation framework to assess the capability of LLMs to function as GPs. Based on this framework, we introduce a general practice benchmark (GPBench), whose data are meticulously annotated by domain experts in accordance with routine clinical practice standards. We evaluate ten state-of-the-art LLMs and analyze their competencies. Our findings indicate that current LLMs are not suitable for autonomous deployment in clinical general practice and that all realistic applications require continuous human oversight; further optimization specifically tailored to the daily responsibilities of GPs remains essential.

大模型评估医疗AI临床推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。