arXiv:2605.03832cs.LG2026-05

构建医疗时间序列跨区域迁移的持续学习基准,解决模型泛化难题。

A Domain Incremental Continual Learning Benchmark for ICU Time Series Model Transportability

论文配图:A Domain Incremental Continual Learning Benchmark for ICU Time Series Model Transportability
图 1 · 摘自论文原文
  • 将跨区域模型迁移建模为领域增量学习问题
  • 验证数据回放与弹性权重固化方法在新地区表现
  • 适合关注临床模型可移植性的研究者与医疗机构

近年来,机器学习在临床结局预测中取得显著进展,但训练模型所需的数据收集、标注和算力资源限制了中小型医院自主开发模型的可行性。一种替代方案是将大型医院训练的模型迁移到小医院,通过本地患者数据微调。然而,现有模型多在单一医院数据上训练与验证,泛化能力存疑。本研究发现美国不同地区测量数据分布与频率存在显著差异。为此,我们提出一个基准,评估模型从源域向全国不同地区迁移的能力。该基准将模型迁移视为领域增量学习任务:尽管预测目标不变,输入数据分布随地域变化,需模型有效适应分布偏移并保留原域关键特征。我们采用数据回放与弹性权重固化(EWC)两种主流领域增量学习方法进行评估。

原文摘要 · Abstract (English)

In recent years, machine learning has made significant progress in clinical outcome prediction, demonstrating increasingly accurate results. However, the substantial resources required for hospitals to train these models, such as data collection, labeling, and computational power, limit the feasibility for smaller hospitals to develop their own models. An alternative approach involves transferring a machine learning model trained by a large hospital to smaller hospitals, allowing them to fine-tune the model on their specific patient data. However, these models are often trained and validated on data from a single hospital, raising concerns about their generalizability to new data. Our research shows that there are notable differences in measurement distributions and frequencies across various regions in the United States. To address this, we propose a benchmark that tests a machine learning model's ability to transfer from a source domain to different regions across the country. This benchmark assesses a model's capacity to learn meaningful information about each new domain while retaining key features from the original domain. Using this benchmark, we frame the transfer of a machine learning model from one region to another as a domain incremental learning problem. While the task of patient outcome prediction remains the same, the input data distribution varies, necessitating a model that can effectively manage these shifts. We evaluate two popular domain incremental learning methods: data replay, which stores examples from previous data sources for fine-tuning on the current source, and Elastic Weight Consolidation (EWC), a model parameter regularization method that maintains features important for both data sources.

持续学习医疗AI模型迁移时间序列

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。