用仿真模型生成灾难场景下的病历数据,检验医疗AI在极端情况下的可靠性。
Generating synthetic electronic health record data using agent-based models to evaluate machine learning robustness under mass casualty incidents

- 构建急诊科代理模型,模拟灾难时患者流量与资源变化。
- 生成的合成数据中,模型预测留院时间的召回率下降18%以上。
- 适合医疗AI开发者和临床系统评估人员参考。
医疗机器学习模型通常基于整理好的真实电子健康记录(EHR)数据进行评估。这类评估的一大局限是无法检验模型在实际部署时面对数据变化的鲁棒性,因为用于模型开发的真实EHR数据难以涵盖所有可能的变化。由灾难引发的大规模伤亡事件(MCIs)正是此类问题的典型场景,会带来罕见、不确定且新颖的数据变化。由于真实世界中MCIs的EHR数据往往稀缺或不可得,部署前评估模型在这些条件下的表现仍具挑战。本文提出一种基于代理的建模方法,生成用于评估医疗机器学习模型鲁棒性的合成EHR数据。我们利用真实EHR数据构建并校准一个急诊科(ED)的代理模型(ABM),显式模拟患者到达、资源容量和临床工作流程。通过调整这些系统条件以反映合理的MCIs情景,该模型生成具有系统行为偏移的合成真实EHR数据。利用这些数据,我们测试了预测住院时长的机器学习模型。结果显示,在MCIs条件下,模型的召回率相比基线系统条件持续下降,导致被遗漏的长期住院患者数量增加。结果表明,系统条件变化对患者结局、EHR数据和模型性能均有显著影响。本研究确立了基于代理模型的合成EHR数据生成是一种主动且系统的评估方法,可用于评估医疗模型在真实数据未覆盖的异常系统状态(如MCIs)下的鲁棒性,支持医疗系统中机器学习模型更安全、有效的部署。
原文摘要 · Abstract (English)
ML models in healthcare are typically evaluated using curated real-world EHR data. A key limitation of such evaluations is that they may fail to assess the robustness of ML models to changes in the data at deployment, which is a common issue because EHR data used for ML model development cannot capture all such changes. Mass casualty incidents (MCIs) caused by disasters are critical instances where this will be an issue, as they induce rare, uncertain, and novel changes to routine system conditions. Because real-world EHR data from MCIs are often limited or unavailable, assessing ML robustness under such conditions before deployment remains challenging. Here, we propose an agent-based modelling approach for generating synthetic EHR data to evaluate the robustness of ML models under MCI scenarios. We use real-world EHR data to develop and calibrate an agent-based model (ABM) of an emergency department (ED) that explicitly models patient arrivals, resource capacity, and clinical workflow. By changing these system conditions to reflect plausible MCI scenarios, the ED model generates synthetic versions of the real-world EHR data that exhibit shifts in system behaviour. Using these synthetic data, we test ML models for predicting length of stay. We observed consistent declines in recall under MCI conditions relative to baseline system conditions, resulting in an increase in the number of patients with prolonged length of stay that were missed by the ML models. These results highlight the impact of changes in system conditions on patient outcomes, EHR data, and ML model performance. Our work establishes ABM-based synthetic EHR data generation as a proactive and systematic approach for evaluating the robustness of ML models under MCI or other system conditions not captured in real-world EHR data, supporting the safer and more effective deployment of ML models in healthcare systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。