用大模型发现电子病历中21.7%患者未服药,导致预测偏差和治疗效果误判。
Revealing Treatment Non-Adherence Bias in Clinical Machine Learning Using Large Language Models
- 用大模型从病历文本提取用药依从性信息
- 发现非依从率21.7%,可使治疗效果估计反转
- 揭示弱势群体在模型误差中被放大
基于3,623名高血压患者的电子健康记录(EHR),我们研究了用药不依从如何引入隐性偏差,从而从根本上扭曲因果推断与预测建模。通过大语言模型(LLM)从临床笔记中提取依从性信息,识别出786名(21.7%)用药不依从患者。我们进一步发现与依从性相关的关键人口统计学和临床因素,以及患者自述的常见原因,如副作用和取药困难。研究显示,该隐性偏差不仅可反转治疗效果估计,还导致模型性能下降最高达5%,且对弱势群体的决策结果和模型误差率产生加剧影响。这凸显了在构建负责任、公平的临床机器学习系统时,必须考虑用药依从性问题。
原文摘要 · Abstract (English)
Machine learning systems trained on electronic health records (EHRs) increasingly guide treatment decisions, but their reliability depends on the critical assumption that patients follow the prescribed treatments recorded in EHRs. Using EHR data from 3,623 hypertension patients, we investigate how treatment non-adherence introduces implicit bias that can fundamentally distort both causal inference and predictive modeling. By extracting patient adherence information from clinical notes using a large language model (LLM), we identify 786 patients (21.7%) with medication non-adherence. We further uncover key demographic and clinical factors associated with non-adherence, as well as patient-reported reasons including side effects and difficulties obtaining refills. Our findings demonstrate that this implicit bias can not only reverse estimated treatment effects, but also degrade model performance by up to 5% while disproportionately affecting vulnerable populations by exacerbating disparities in decision outcomes and model error rates. This highlights the importance of accounting for treatment non-adherence in developing responsible and equitable clinical machine learning systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。