arXiv:2506.00095cs.CYcs.AI2025-06被引 1

首个针对肝胆胰疾病的LLM评估基准,揭示现有模型在真实临床诊断中表现不足。

ClinBench-HPB: A Clinical Benchmark for Evaluating LLMs in Hepato-Pancreato-Biliary Diseases

  • 构建包含3535道选择题和337个真实病例的肝胆胰疾病评测集
  • 商业大模型在复杂住院病例上准确率显著下降,医疗专用模型泛化能力有限
  • 适合医学AI研究者、临床辅助系统开发者使用,推动真实场景下的模型评估

肝胆胰(HPB)疾病因高发病率与死亡率构成全球公共卫生挑战。尽管大语言模型(LLMs)在通用医学问答任务中表现优异,但现有评估基准多源自标准化考试或人工设计题目,缺乏对HPB疾病的覆盖与真实临床案例。为此,我们系统构建了一个涵盖国际疾病分类第10版(ICD-10)全部33个主类与465个子类的HPB疾病评估基准,包含3,535道闭合式多选题和337个开放式真实诊断病例。多选题来自公开数据集与合成数据,临床病例则采集自权威医学期刊、病例共享平台及合作医院。在该基准(ClinBench-HPB)上评估商业与开源通用及医疗类LLMs发现:虽商业模型在医学考试题上表现良好,但在复杂住院病例诊断任务中性能明显下降;医疗专用模型对HPB疾病的泛化能力亦有限。结果揭示当前LLMs在肝胆胰疾病领域存在关键局限,凸显未来需聚焦真实复杂临床诊断而非简单考试题目的迫切需求。基准将发布于https://clinbench-hpb.github.io。

原文摘要 · Abstract (English)

Hepato-pancreato-biliary (HPB) disorders represent a global public health challenge due to their high morbidity and mortality. Although large language models (LLMs) have shown promising performance in general medical question-answering tasks, the current evaluation benchmarks are mostly derived from standardized examinations or manually designed questions, lacking HPB coverage and clinical cases. To address these issues, we systematically eatablish an HPB disease evaluation benchmark comprising 3,535 closed-ended multiple-choice questions and 337 open-ended real diagnosis cases, which encompasses all the 33 main categories and 465 subcategories of HPB diseases defined in the International Statistical Classification of Diseases, 10th Revision (ICD-10). The multiple-choice questions are curated from public datasets and synthesized data, and the clinical cases are collected from prestigious medical journals, case-sharing platforms, and collaborating hospitals. By evalauting commercial and open-source general and medical LLMs on our established benchmark, namely ClinBench-HBP, we find that while commercial LLMs perform competently on medical exam questions, they exhibit substantial performance degradation on HPB diagnosis tasks, especially on complex, inpatient clinical cases. Those medical LLMs also show limited generalizability to HPB diseases. Our results reveal the critical limitations of current LLMs in the domain of HPB diseases, underscoring the imperative need for future medical LLMs to handle real, complex clinical diagnostics rather than simple medical exam questions. The benchmark will be released at https://clinbench-hpb.github.io.

医学大模型临床评估肝胆胰疾病

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。