构建中文医疗质控指标计算数据集,评测大模型在真实临床任务中的表现
CMQCIC-Bench: A Chinese Benchmark for Evaluating Large Language Models in Medical Quality Control Indicator Calculation
- 提出基于临床事实的推理分解方法,分离验证与推断步骤
- 在785条真实病历上测试,新方法优于传统思维链
- 适合医疗AI研究者、临床决策系统开发者参考
医疗质量控制指标是评估医疗机构服务资质的关键。随着GPT-4等大语言模型在医疗领域表现出色,利用其进行医疗质量控制指标计算(MQCIC)具有广阔前景。本文首先提出一个真实世界任务MQCIC,构建了基于中文电子病历的开源数据集CMQCIC-Bench,包含785个实例和76项指标。其次,提出一种半自动规则表示增强方法,并设计临床事实驱动的推理规则(CF-IR),将临床事实验证与推理规则拆解分离。在20个代表性大模型上开展全面实验,结果表明CF-IR在MQCIC任务中优于思维链方法。进一步通过错误分析,揭示了临床事实验证与推理能力的差异,为提升性能提供依据。数据集与代码已开源:https://github.com/YuY-2001/C-MQCIC。
原文摘要 · Abstract (English)
Medical quality control indicators are essential to assess the qualifications of healthcare institutions for medical services. With the impressive performance of large language models (LLMs) like GPT-4 in the medical field, leveraging these technologies for the Medical Quality Control Indicator Calculation (MQCIC) presents a promising approach. In this work, (1) we introduce a real-world task MQCIC and propose an open-source Chinese electronic medical records (EMRs)-based dataset (CMQCIC-Bench) comprising 785 instances and 76 indicators. (2) We propose a semi-automatic method to enhance the rule representation. Then we propose the Clinical Facts-based Inferential Rule (CF-IR) method that disentangles the clinical fact verification and inferential rule reasoning actions. (3) We conduct comprehensive experiments on 20 representative LLMs, covering general and medical models. Our findings reveal that CF-IR outperforms Chain-of-Thought methods in MQCIC tasks. (4) We conduct an error analysis and investigate the capabilities of clinical fact verification and inferential rule reasoning, providing insights to improve performance in the MQCIC further. The dataset and code is available in this repository https://github.com/YuY-2001/C-MQCIC.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。