用大模型预测未记录的吸烟状态,纠正偏倚估计超声影响死亡率的因果效应
Controlling for Unobserved Confounding with Large Language Model Classification of Patient Smoking Status
- 用临床文本大模型预测缺失的吸烟状态作为潜在混杂因素
- 在MIMIC数据集上修正分类误差后,得出超声对死亡率的因果影响
- 为真实世界医疗数据中隐藏混杂因子的处理提供新方法
因果推断是循证医学的核心目标。当随机化不可行时,需依赖观察性数据进行回顾性分析。但此类分析常依赖无未观测混杂的假设,而实际中重要变量未被记录时该假设常不成立。先前研究提出用机器学习补全未观测变量并校正分类误差,从而恢复无偏因果估计,但限于合成数据、简单分类器和二元变量。本文扩展该方法,使用基于临床笔记训练的大语言模型预测患者吸烟状态(原本未记录的混杂因子),并在MIMIC数据集中对分类结果应用测量误差校正,估计经胸超声心动图对死亡率的因果效应。
原文摘要 · Abstract (English)
Causal understanding is a fundamental goal of evidence-based medicine. When randomization is impossible, causal inference methods allow the estimation of treatment effects from retrospective analysis of observational data. However, such analyses rely on a number of assumptions, often including that of no unobserved confounding. In many practical settings, this assumption is violated when important variables are not explicitly measured in the clinical record. Prior work has proposed to address unobserved confounding with machine learning by imputing unobserved variables and then correcting for the classifier's mismeasurement. When such a classifier can be trained and the necessary assumptions are met, this method can recover an unbiased estimate of a causal effect. However, such work has been limited to synthetic data, simple classifiers, and binary variables. This paper extends this methodology by using a large language model trained on clinical notes to predict patients' smoking status, which would otherwise be an unobserved confounder. We then apply a measurement error correction on the categorical predicted smoking status to estimate the causal effect of transthoracic echocardiography on mortality in the MIMIC dataset.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。