用GPT-4o自动识别电子病历中的认知障碍阶段,准确率超0.83。
Evaluating GPT's Capability in Identifying Stages of Cognitive Impairment from Electronic Health Data
- 零样本调用GPT-4o分析医生笔记判断认知阶段
- 在769例患者中达成0.83的加权卡帕系数
- 适用于临床辅助诊断与大规模研究数据构建
从电子健康记录(EHR)中识别认知障碍对及时诊断和研究均至关重要。相关线索常隐藏于未结构化的医生笔记中,但人工查阅费时且易出错。本研究评估了零样本GPT-4o在两项任务中的表现:首先,在麻省总医院记忆门诊的769名患者专科学习笔记上,判断整体临床痴呆评分(CDR),获得加权卡帕系数0.83;其次,在860名医保患者的三年内所有笔记上,区分正常认知、轻度认知障碍(MCI)和痴呆,相比专家审阅,加权卡帕系数达0.91,高信心案例下达0.96。结果表明,GPT-4o具备作为可扩展病历审查工具的潜力,未来可用于研究数据生成与临床辅助诊断。
原文摘要 · Abstract (English)
Identifying cognitive impairment within electronic health records (EHRs) is crucial not only for timely diagnoses but also for facilitating research. Information about cognitive impairment often exists within unstructured clinician notes in EHRs, but manual chart reviews are both time-consuming and error-prone. To address this issue, our study evaluates an automated approach using zero-shot GPT-4o to determine stage of cognitive impairment in two different tasks. First, we evaluated the ability of GPT-4o to determine the global Clinical Dementia Rating (CDR) on specialist notes from 769 patients who visited the memory clinic at Massachusetts General Hospital (MGH), and achieved a weighted kappa score of 0.83. Second, we assessed GPT-4o's ability to differentiate between normal cognition, mild cognitive impairment (MCI), and dementia on all notes in a 3-year window from 860 Medicare patients. GPT-4o attained a weighted kappa score of 0.91 in comparison to specialist chart reviews and 0.96 on cases that the clinical adjudicators rated with high confidence. Our findings demonstrate GPT-4o's potential as a scalable chart review tool for creating research datasets and assisting diagnosis in clinical settings in the future.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。