用专家制定的评分标准引导大模型,从病历中精准识别阿片类药物滥用患者。
A Rubric-Guided Large Language Model Solution for Opioid Use Disorder Computable Phenotyping
- 基于18项专家评分标准,指导大模型提取病历中的关键证据。
- 在253例患者上达到0.774 F1和0.934 AUROC,优于传统方法。
- 结果可解释性强,适合临床研究与医疗决策支持场景。
阿片类药物滥用(OUD)是美国严重的公共卫生问题,但因其诊断编码缺失且证据深藏于临床文本中,难以从电子健康记录(EHR)中准确识别。本研究提出一种基于评分标准引导的大语言模型(LLM)框架,结合优化提示技术(OPRO),用于OUD可计算表型(CP)识别。该框架采用18项由专家制定的评分标准,指导LLM自动提取支持性文本以判断是否为OUD阳性。两位佛罗里达大学医生(GMR 和 WMG)对253名患者进行人工病历审查,其中68例为OUD阳性。所提方法在测试中获得0.774的最高F1得分和0.934的AUROC,相比基于机器学习的CP方法和零样本大模型,相对F1提升分别达12.8%和44.4%。该方法能将大模型提取的证据与表型判断关联,显著提升可解释性。
原文摘要 · Abstract (English)
Opioid use disorder (OUD) remains a public health crisis in the United States, yet it is difficult to identify from electronic health records (EHRs) because missing diagnosis codes and supporting evidence are buried in clinical narratives. Accurate OUD identification is critical to support interventions and improve health outcomes. This study developed a rubric-guided large language model (LLM) that incorporated Optimization by PROmpting (OPRO) for OUD computable phenotyping (CP). The framework used an 18-item, expert-identified rubric to instruct LLMs to automatically extract critical text with supporting evidence to determine OUD flags. Two UF Health physicians (GMR and WMG) chart-reviewed 253 patients, including 68 OUD-positive cases. Our LLM-based computable phenotype (CP) achieved the best F1 score of 0.774 and an AUROC of 0.934, outperforming the machine learning-based CP using EHR and natural language processing-extracted variables, and zero-shot LLMs by relative F1 improvements of 12.8% and 44.4%, respectively. The proposed LLM-based CP could link LLM-extracted evidence to OUD phenotyping for better explainability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。