arXiv:2410.03502cs.CL2024-10EMNLP被引 21

构建中文临床大模型评估基准,测试真实医疗场景下的表现

CliMedBench: A Large-Scale Chinese Benchmark for Evaluating Medical Large Language Models in Clinical Scenarios

论文配图:CliMedBench: A Large-Scale Chinese Benchmark for Evaluating Medical Large Language Models in Clinical Scenarios
图 1 · 摘自论文原文
  • 设计14个临床核心场景,覆盖7个关键维度的3.37万道真实医学题
  • 中文医疗大模型在推理与事实一致性上表现不足,通用模型潜力大
  • 适合医疗AI研究者、临床决策系统开发者参考

随着大语言模型在多个领域的发展,临床医疗场景亟需统一的评估标准,要求模型进行严格检验。我们提出CliMedBench,一个涵盖14个专家指导的核心临床场景的综合性基准,专门用于评估大语言模型在7个核心维度上的医学能力。该基准包含33,735道问题,源自顶级三甲医院的真实病历报告和真实考试题。通过多项验证,其可靠性已得到确认。对现有大模型的实验显示:(i) 中文医疗大模型在此基准上表现不佳,尤其在医学推理与事实一致性方面,凸显临床知识与诊断准确性的提升需求;(ii) 多个通用领域大模型展现出显著潜力,但许多医疗专用模型因输入容量有限,难以实际应用。这些发现揭示了大模型在临床场景中的优劣,为医学研究提供了关键洞见。

原文摘要 · Abstract (English)

With the proliferation of Large Language Models (LLMs) in diverse domains, there is a particular need for unified evaluation standards in clinical medical scenarios, where models need to be examined very thoroughly. We present CliMedBench, a comprehensive benchmark with 14 expert-guided core clinical scenarios specifically designed to assess the medical ability of LLMs across 7 pivot dimensions. It comprises 33,735 questions derived from real-world medical reports of top-tier tertiary hospitals and authentic examination exercises. The reliability of this benchmark has been confirmed in several ways. Subsequent experiments with existing LLMs have led to the following findings: (i) Chinese medical LLMs underperform on this benchmark, especially where medical reasoning and factual consistency are vital, underscoring the need for advances in clinical knowledge and diagnostic accuracy. (ii) Several general-domain LLMs demonstrate substantial potential in medical clinics, while the limited input capacity of many medical LLMs hinders their practical use. These findings reveal both the strengths and limitations of LLMs in clinical scenarios and offer critical insights for medical research.

医疗大模型评测基准临床推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。