解决医学个性化治疗中偏差与精度的矛盾,提升重症患者预测准确率。
Resolving the bias-precision paradox with stochastic causal representation learning for personalized medicine

- 用随机匹配替代全局对抗平衡,缓解混杂偏差
- 在两个大型ICU队列中误差降低11.5%,高风险任务召回率显著提升
- 可解释性强,适合临床实时决策支持,优于医生和大模型
从纵向观察数据中估计个体化治疗效应是数据驱动医学的核心,但现有方法存在根本局限:减少混杂偏差常会抑制临床相关的异质性,导致患者特异性预测性能下降。本文将这一矛盾定义为因果表示学习中的偏差-精度悖论,并提出基于采样的最大均值差异(sMMD)方法,以子集级匹配替代全局对抗平衡。该方法被应用于具有归因基础可解释性的反事实结果预测框架。在两个大规模ICU队列(n = 27,783)上,该框架在分布偏移条件下提升了预测准确性,误差最多降低11.5%,高风险任务召回率大幅提高。机制分析显示,sMMD能选择性保留临床关键变量。人机评估表明,该方法优于医学生训练组和大语言模型,在提升医生诊断准确率14.7%的同时缩短决策时间,实现可解释、实时的临床决策支持。
原文摘要 · Abstract (English)
Estimating individualized treatment effects from longitudinal observational data is central to data-driven medicine, yet existing methods face a fundamental limitation: reducing confounding bias often suppresses clinically informative heterogeneity, degrading patient-specific predictions. Here, we identify this tension as a bias-precision paradox in causal representation learning and introduce sampling-based maximum mean discrepancy (sMMD), a stochastic alignment strategy that replaces global adversarial balancing with subset-level matching. We instantiate this approach in a framework for counterfactual outcome prediction with attribution-grounded interpretability. Across two large-scale ICU cohorts (n = 27,783), our framework improves accuracy under distribution shift, reducing error by up to 11.5% and substantially increasing recall in high-risk tasks. Mechanistic analyses show that sMMD selectively preserves clinically decisive variables. In human-AI evaluation, our method outperforms clinicians-in-training and large language models, and improves clinician accuracy by 14.7% while reducing decision time, enabling interpretable, real-time clinical decision support.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。