首个评估医学笔记错误检测与修正能力的基准,检验大模型医疗推理水平。
MEDEC: A Benchmark for Medical Error Detection and Correction in Clinical Notes
- 构建涵盖五类错误的3848条临床文本数据集,含未见过的真实病历。
- 大模型在纠错任务中表现良好但仍逊于专业医生,尤其在复杂推理上。
- 适合评估医疗AI可靠性,推动临床生成内容的质量保障研究。
尽管大型语言模型(LLMs)在多项医学考试中表现优于平均人类水平,但尚无研究系统评估其验证或修正现有/生成医学文本正确性与一致性的能力。本文提出MEDEC(https://github.com/abachaa/MEDEC),首个公开可用的临床笔记错误检测与修正基准,涵盖诊断、管理、治疗、药物治疗和致病生物体五类错误。该数据集包含3,848条临床文本,其中488条来自三家美国医院系统且未被任何大模型见过。该数据集已用于MEDIQA-CORR共享任务,评估17个参与系统的表现。我们描述了数据构建方法,并评估了近期主流模型(如o1-preview、GPT-4、Claude 3.5 Sonnet、Gemini 2.0 Flash)在需医学知识与推理的任务中的表现。同时,两名医学博士在相同测试集上完成任务作为对照。结果表明,MEDEC具备足够挑战性,可有效评估模型验证与修正医疗文本的能力。尽管当前大模型在错误检测与修正上表现良好,但仍不及专业医生。我们分析了差距成因、实验洞察、现有评估指标局限,并为未来研究提供方向。
原文摘要 · Abstract (English)
Several studies showed that Large Language Models (LLMs) can answer medical questions correctly, even outperforming the average human score in some medical exams. However, to our knowledge, no study has been conducted to assess the ability of language models to validate existing or generated medical text for correctness and consistency. In this paper, we introduce MEDEC (https://github.com/abachaa/MEDEC), the first publicly available benchmark for medical error detection and correction in clinical notes, covering five types of errors (Diagnosis, Management, Treatment, Pharmacotherapy, and Causal Organism). MEDEC consists of 3,848 clinical texts, including 488 clinical notes from three US hospital systems that were not previously seen by any LLM. The dataset has been used for the MEDIQA-CORR shared task to evaluate seventeen participating systems [Ben Abacha et al., 2024]. In this paper, we describe the data creation methods and we evaluate recent LLMs (e.g., o1-preview, GPT-4, Claude 3.5 Sonnet, and Gemini 2.0 Flash) for the tasks of detecting and correcting medical errors requiring both medical knowledge and reasoning capabilities. We also conducted a comparative study where two medical doctors performed the same task on the MEDEC test set. The results showed that MEDEC is a sufficiently challenging benchmark to assess the ability of models to validate existing or generated notes and to correct medical errors. We also found that although recent LLMs have a good performance in error detection and correction, they are still outperformed by medical doctors in these tasks. We discuss the potential factors behind this gap, the insights from our experiments, the limitations of current evaluation metrics, and share potential pointers for future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。