直接从不完整病历中学习表示,提升临床任务表现
Learning Representations from Incomplete EHR Data with Dual-Masked Autoencoding
- 结合原始缺失标记与人为遮蔽观测值,双掩码重建
- 在两个数据集上多项临床任务性能超越基线
- 无需补全缺失值,也能学出有临床意义的表征
电子健康记录(EHR)本就带有缺失。临床医生选择性地开具检查,导致每张患者表仅包含部分生理状态信息。以往方法或先补全数据,或用特殊标记表示缺失,或只优化补全效果,限制了下游任务表示能力并把所有未观测项带入编码器。我们提出AID-MAE——一种增强-内在双掩码自编码器,通过融合记录固有的缺失标记与人为遮蔽部分观测值,在预训练中直接从不完整表中学习。两类被遮蔽条目均不进入编码器,注意力仅作用于实际观测内容。AID-MAE在两个数据集上的多个临床任务中持续优于强基线。实验表明,恢复缺失值并非学习前提,所学表征无需监督即可保留临床结构。
原文摘要 · Abstract (English)
Electronic health records (EHR) arrive masked. Clinicians order measurements selectively, and any patient table thus contains only a subset of the values that characterize the underlying physiological state. Prior masked modeling approaches on EHR data either impute the table before learning, represent missingness through a dedicated placeholder signal, or optimize solely for imputation, which limits the representations they learn for downstream clinical tasks and carries every unobserved entry through the encoder. We introduce AID-MAE, an Augmented-Intrinsic Dual-Masked Autoencoder that learns directly from incomplete tables by combining the intrinsic mask the record already carries with an augmented mask that hides a subset of observed values for reconstruction during pretraining. Neither type of masked entry enters the encoder, so attention operates only over what was observed. AID-MAE achieves consistent improvements over strong baselines across multiple clinical tasks on two datasets. Across experiments, we discuss that recovering the missing entries is not a prerequisite for learning and show that the representations learned carry clinical structure without supervision.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。