用自监督方法提升医疗数据缺失值填补的准确性与公平性
Representation Learning of Lab Values via Masked AutoEncoders
- 基于Transformer的掩码自编码框架,联合建模检验值与时间戳
- 在MIMIC-IV上优于多种基线方法,尤其在不同人群间表现均衡
- 适合关注临床预测公平性与高精度数据修复的研究者
电子健康记录中实验室数值的准确补全对实现稳健的临床预测和减少医疗AI中的偏差至关重要。现有方法如XGBoost、softimpute、GAIN、期望最大化(EM)和MICE难以捕捉EHR数据中复杂的时序与上下文依赖关系,尤其在低代表群体中表现不佳。本文提出Lab-MAE,一种基于Transformer的掩码自编码框架,利用自监督学习进行连续序列化验值的补全。Lab-MAE引入结构化编码机制,联合建模检验项目值及其对应时间戳,显式捕捉时序依赖。在MIMIC-IV数据集上的实证评估表明,Lab-MAE在均方根误差(RMSE)、决定系数(R2)和沃斯泰因距离(WD)等多项指标上显著优于XGBoost、softimpute、GAIN、EM和MICE等先进基线方法。值得注意的是,Lab-MAE在不同患者人口统计学群体间表现出均衡性能,推动了临床预测的公平性。我们进一步探究了随访检验值作为潜在捷径特征的作用,发现当此类数据不可用时,Lab-MAE仍具鲁棒性。研究结果表明,针对EHR数据特性优化的Transformer架构为更精准、更公平的临床补全提供了基础模型。此外,我们还测量并比较了Lab-MAE与XGBoost模型的碳足迹,凸显其环境成本。
原文摘要 · Abstract (English)
Accurate imputation of missing laboratory values in electronic health records (EHRs) is critical to enable robust clinical predictions and reduce biases in AI systems in healthcare. Existing methods, such as XGBoost, softimpute, GAIN, Expectation Maximization (EM), and MICE, struggle to model the complex temporal and contextual dependencies in EHR data, particularly in underrepresented groups. In this work, we propose Lab-MAE, a novel transformer-based masked autoencoder framework that leverages self-supervised learning for the imputation of continuous sequential lab values. Lab-MAE introduces a structured encoding scheme that jointly models laboratory test values and their corresponding timestamps, enabling explicit capturing temporal dependencies. Empirical evaluation on the MIMIC-IV dataset demonstrates that Lab-MAE significantly outperforms state-of-the-art baselines such as XGBoost, softimpute, GAIN, EM, and MICE across multiple metrics, including root mean square error (RMSE), R-squared (R2), and Wasserstein distance (WD). Notably, Lab-MAE achieves equitable performance across demographic groups of patients, advancing fairness in clinical predictions. We further investigate the role of follow-up laboratory values as potential shortcut features, revealing Lab-MAE's robustness in scenarios where such data is unavailable. The findings suggest that our transformer-based architecture, adapted to the characteristics of EHR data, offers a foundation model for more accurate and fair clinical imputation. In addition, we measure and compare the carbon footprint of Lab-MAE with the a XGBoost model, highlighting its environmental requirements.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。