用因果隐马尔可夫模型纠正医疗诊断偏倚,让预测更公平。
Correcting heterogeneous diagnostic bias when developing clinical prediction models using causal hidden Markov models
- 构建潜变量马尔可夫模型,模拟疾病进展与检测延迟过程。
- 在模拟和真实数据中,将误诊率偏差从1.34降至1.02。
- 适合关注医疗公平性、模型校准的临床研究者使用。
常规诊疗中,高风险人群通常接受更频繁的检测,而性别、种族等属性也影响检测频率,导致不同群体诊断率不均,引发标签错误与系统性模型偏差。本文提出一种基于因果推断的隐马尔可夫模型方法,目标是估计个体在诊断率与参照组一致的反事实场景下的诊断概率。通过建模纵向疾病进展过程,将确诊结果视为隐藏疾病阶段的观测输出。在模拟数据中,该方法将未诊断群体的观察:预期比从1.34(标准差0.09)降至1.02(0.09),显著改善整体校准度;即使假设不成立,仍优于标准模型。在慢性肾病预测案例中,发现糖尿病使6个月尿白蛋白肌酐比值检测率提升10.36倍(95%置信区间9.80–11.02),通过修正非糖尿病患者反事实诊断率,使模型观察:预期比从1.55(1.51–1.59)降至1.01(0.98–1.04)。
原文摘要 · Abstract (English)
In routine care, individuals identified a priori as high-risk are usually tested for conditions more frequently. Protected attributes, such as sex or ethnicity may also determine testing frequency. Such heterogeneous detection rates across a population induce label error. This causes systematic model error for specific groups and biases performance metrics during validation. This paper proposes a method to correct for such bias in prediction models due to differential diagnostic delay. We use a causal inference framework to define our target estimand: an individual's diagnosis probability in a counterfactual scenario where their diagnosis rate matches that of a reference group. We model the longitudinal process as a hidden Markov model, in which confirmatory test results are emissions from a latent progressive disease stage. We validate our approach in simulated data and apply it to a case study of chronic kidney disease prediction using electronic health records. In simulations, our method reduces prediction bias and improves calibration-in-the-large, correcting the Observed:Expected ratio in the underdiagnosed group from 1.34 (standard deviation: 0.09) in a model developed without any correction for underdiagnosis bias to 1.02 (0.09). Violations of assumptions in the simulation affected the estimation of model parameters, but the proposed approach nonetheless remained better calibrated than the standard model. In the clinical case study, we identify diabetes as the main driver of observability, with an odds ratio of 10.36 (95% confidence interval, 9.80 - 11.02) in 6-month urine albumin-creatinine ratio testing rate. Using our approach to predict the counterfactual diagnostic rate in patients without diabetes, we improved the Observed:Expected ratio of a developed clinical prediction model from 1.55 (1.51 - 1.59) to 1.01 (0.98 - 1.04).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。