解决医疗模型跨国家部署时关键数据缺失问题,提升预测可靠性。
Distributionally Robust Transfer Learning with Structurally Missing Covariates, with Application to Cross-National Cardiac Arrest Prediction

- 将变量分为共享与缺失两类,不补全缺失数据
- 在未知目标分布下优化最差情况下的预测性能
- 适合缺乏标签和完整数据的跨国医疗预测场景
在不同医疗系统间部署临床预测模型时,常因关键训练变量在目标域不可用且标注数据有限而失败。例如,院外心脏骤停(OHCA)高精度模型依赖于高资源地区常规采集的详细院前数据,但在许多国际注册数据库中无法获取。现有方法或丢弃缺失变量(损失信息),或依赖无法验证的目标分布假设。本文提出DRUM(分布鲁棒无监督迁移学习,针对结构缺失协变量),可在某些协变量完全缺失且无标签的情况下,将模型迁移到目标人群。DRUM将协变量分为跨域共享部分(X)和仅源域观测的缺失部分(A)。不进行缺失变量插补,而是通过神经网络生成器,在未知目标条件分布 $A ackslashmid X$ 上优化最坏情况下的预测性能,并由鲁棒性参数控制与源域条件的允许偏差。进一步设计偏差校正程序以降低对无关估计误差的敏感性。模拟实验显示,在分布偏移下均值和最差情况预测误差均有显著改善。应用于从美国注册库向多个亚洲注册库迁移OHCA预测模型,尽管院前变量未记录,DRUM仍实现更优校准性和跨站点临床分类性能。
原文摘要 · Abstract (English)
Deploying clinical prediction models across healthcare systems often fails when key training covariates are unavailable at deployment and labeled outcomes are limited in the target domain. For example, high-performing models for out-of-hospital cardiac arrest (OHCA) rely on detailed prehospital measurements routinely collected in high-resource settings but unavailable in many international registries. Existing methods either discard missing covariates, sacrificing predictive information, or rely on untestable assumptions about their target distribution. We propose DRUM (\underline{D}istributionally \underline{R}obust \underline{U}nsupervised transfer learning with structurally \underline{M}issing covariates), a framework that transfers prediction models to target populations where certain covariates are structurally absent and outcome labels are unavailable. DRUM partitions covariates into shared components ($X$), observed across all settings, and missing components ($A$), observed only in the source. Rather than imputing missing covariates, DRUM optimizes worst-case predictive performance over the unknown target distribution of $A \mid X$ using a neural network generator, with a robustness parameter controlling allowable deviation from the source conditional. We further develop a bias correction procedure that reduces sensitivity to nuisance estimation error. Simulations show substantial improvements in both mean and worst-case prediction error under distribution shift. Applied to cross-national OHCA prediction, transferring models from a US registry to multiple Asian registries where prehospital variables are unrecorded, DRUM yields better-calibrated predictions and improved clinical classification performance across sites.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。