用专业评分标准验证大模型生成病历质量,发现与真人接近但略逊一筹。
Assessing the Quality of AI-Generated Clinical Notes: A Validated Evaluation of a Large Language Model Scribe
- 用医生专家评分工具PDQI9对比大模型与真人书写的病历质量。
- 真人病历平均得分4.25/5,大模型为4.20/5,差异显著但不大。
- 该方法可作为评估AI病历质量的实用工具,适合医疗AI开发者参考。
美国多地医疗机构已开始使用生成式人工智能(AI)工具充当医生记录员,以减轻临床文书负担。然而,目前尚无统一方法评估此类AI记录的质量。为此,我们开展了一项盲法研究,比较大型语言模型(LLM)生成的病历与领域专家撰写病历在音频记录的临床就诊中的表现。采用医师文档质量工具(PDQI9)的量化指标作为评估框架,覆盖五个医学专科的临床专家对97例患者就诊的「金标准」病历(Gold notes)和「环境模式」病历(Ambient notes)进行评分。每专科两名评估者独立打分。结果显示,全科、骨科及妇产科的评估者间一致性较高(RWG > 0.7),儿科和心脏病学中为中等到高一致(RWG 0.5–0.7)。总体上,金标准病历得分为4.25/5,环境模式病历得分为4.20/5(p = 0.04),差异虽小但具统计意义。结果表明PDQI9可用于实际评估大模型生成病历质量。
原文摘要 · Abstract (English)
In medical practices across the United States, physicians have begun implementing generative artificial intelligence (AI) tools to perform the function of scribes in order to reduce the burden of documenting clinical encounters. Despite their widespread use, no established methods exist to gauge the quality of AI scribes. To address this gap, we developed a blinded study comparing the relative performance of large language model (LLM) generated clinical notes with those from field experts based on audio-recorded clinical encounters. Quantitative metrics from the Physician Documentation Quality Instrument (PDQI9) provided a framework to measure note quality, which we adapted to assess relative performance of AI generated notes. Clinical experts spanning 5 medical specialties used the PDQI9 tool to evaluate specialist-drafted Gold notes and LLM authored Ambient notes. Two evaluators from each specialty scored notes drafted from a total of 97 patient visits. We found uniformly high inter rater agreement (RWG greater than 0.7) between evaluators in general medicine, orthopedics, and obstetrics and gynecology, and moderate (RWG 0.5 to 0.7) to high inter rater agreement in pediatrics and cardiology. We found a modest yet significant difference in the overall note quality, wherein Gold notes achieved a score of 4.25 out of 5 and Ambient notes scored 4.20 out of 5 (p = 0.04). Our findings support the use of the PDQI9 instrument as a practical method to gauge the quality of LLM authored notes, as compared to human-authored notes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。