根据检验指标波动性动态掩码,提升电子病历模型表现
Coefficient of Variation Masking: A Volatility-Aware Strategy for EHR Foundation Models
- 按指标波动程度自适应调整掩码概率,更贴合临床实际
- 在大型检验数据集上重建精度与下游预测性能显著提升
- 适合需要捕捉急性病情变化的临床预训练模型研究者
掩码自编码器(MAEs)被广泛用于电子健康记录(EHR)以学习通用表示,支持多样临床任务。然而,现有方法通常采用均匀随机掩码,隐含假设所有特征可预测性相同。实际上,实验室检验指标存在显著波动性差异:部分生物标志物(如钠离子)稳定,而另一些(如乳酸)波动剧烈,更难建模。临床上,高波动性指标常提示急性病理状态,需更复杂建模以捕捉其时间模式。本文提出波动性感知的预训练策略——变异系数掩码(CV-Masking),根据各特征内在变异性自适应调整掩码概率。结合仅基于数值的掩码目标,该方法在多个基准上优于随机和方差基策略。大规模实验室检验数据实验表明,CV-Masking显著提升重建效果、下游预测性能并加速收敛,生成更具鲁棒性和临床意义的EHR表示。
原文摘要 · Abstract (English)
Masked autoencoders (MAEs) are increasingly applied to electronic health records (EHR) for learning general-purpose representations that support diverse clinical tasks. However, existing approaches typically rely on uniform random masking, implicitly assuming all features are equally predictable. In reality, laboratory tests exhibit substantial heterogeneity in volatility: some biomarkers (e.g., sodium) remain stable, while others (e.g., lactate) fluctuate considerably and are more difficult to model. Clinically, volatile biomarkers often signal acute pathophysiology and require more sophisticated modeling to capture their complex temporal patterns. We propose a volatility-aware pretraining strategy, Coefficient of Variation Masking (CV-Masking), that adaptively adjusts masking probabilities according to the intrinsic variability of each feature. Combined with a value-only masking objective aligned with clinical workflows, CV-Masking yields systematic improvements over random and variance-based strategies. Experiments on a large panel of laboratory tests show that CV-Masking enhances reconstruction, improves downstream predictive performance, and accelerates convergence, producing more robust and clinically meaningful EHR representations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。