arXiv:2507.21783stat.APcs.LG2025-07被引 4

用锚回归提升重症监护模型跨医院泛化能力

Domain Generalization and Adaptation in Intensive Care with Anchor Regression

  • 基于锚回归与树形扩展方法,建模多中心重症数据分布差异
  • 在9个数据库40万患者上,外推性能显著提升,尤其对差异大的医院
  • 提出三阶段框架,指导外部数据使用策略,适合临床部署研究者

临床预测模型在新医院部署时常因分布偏移而性能下降。本文在包含9个不同重症监护室(ICU)数据库、共40万患者的大型数据集上,开展因果启发的域泛化大规模研究。采用锚回归并引入一种新型树基非线性扩展——锚增强,结果表明锚正则化显著提升了跨域性能,尤其在目标域差异最大时效果更明显。该方法对理论假设(如锚外生性)的违反也表现出鲁棒性。此外,提出一个新概念框架,通过评估目标域可用数据量下的性能表现,识别出三种情形:(i) 域泛化阶段,仅使用外部模型即可;(ii) 域适应阶段,微调外部模型最优;(iii) 数据充裕阶段,外部数据不再带来增益。

原文摘要 · Abstract (English)

The performance of predictive models in clinical settings often degrades when deployed in new hospitals due to distribution shifts. This paper presents a large-scale study of causality-inspired domain generalization on heterogeneous multi-center intensive care unit (ICU) data. We apply anchor regression and introduce anchor boosting, a novel, tree-based nonlinear extension, to a large dataset comprising 400,000 patients from nine distinct ICU databases. We find that anchor regularization yields improvements of out-of-distribution performance, particularly for the most dissimilar target domains. The methods appear robust to violations of theoretical assumptions, such as anchor exogeneity. Furthermore, we propose a novel conceptual framework to quantify the utility of large external data datasets. By evaluating performance as a function of available target-domain data, we identify three regimes: (i) a domain generalization regime, where only the external model should be used, (ii) a domain adaptation regime, where refitting the external model is optimal, and (iii) a data-rich regime, where external data provides no additional value.

域泛化重症监护锚回归临床建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。