用半监督学习从海量病历数据中提取更精准的患者表型嵌入。
Semi-supervised Clustering Through Representation Learning of Large-scale EHR Data
- 基于泊松混合模型与预训练编码嵌入,构建多疾病领域患者表征。
- 在部分标注数据下,相比现有方法显著提升对多发性硬化症残疾预测的准确性。
- 适合处理高维稀疏病历数据,尤其适用于标签稀缺的临床预测任务。
电子健康记录(EHR)为个性化医疗提供了丰富的现实世界数据,可揭示疾病进展、治疗反应和患者预后。然而,其稀疏性、异质性和高维度使得建模困难,且缺乏标准化真实标签进一步阻碍了预测模型的发展。为此,我们提出SCORE——一种半监督表示学习框架,通过患者嵌入捕捉多领域疾病特征。SCORE采用泊松自适应潜在因子混合(PALM)模型,结合预训练代码嵌入来刻画编码特征,并提取有意义的患者表型与嵌入。为应对大规模数据的计算挑战,引入混合期望最大化(EM)与高斯变分近似(GVA)算法,利用有限标注数据优化大量未标注样本的估计。我们理论上建立了该混合方法的收敛性,量化了GVA误差,并推导出在发散嵌入维度下的SCORE误差率。分析表明,融入未标注数据可提高精度并降低对标签稀缺的敏感性。大量模拟实验验证了SCORE在有限样本下优于现有方法。最后,我们将SCORE应用于多发性硬化症(MS)患者残疾状态预测,使用部分标注的EHR数据,结果表明其生成的患者嵌入更具信息量且预测性能更优。
原文摘要 · Abstract (English)
Electronic Health Records (EHR) offer rich real-world data for personalized medicine, providing insights into disease progression, treatment responses, and patient outcomes. However, their sparsity, heterogeneity, and high dimensionality make them difficult to model, while the lack of standardized ground truth further complicates predictive modeling. To address these challenges, we propose SCORE, a semi-supervised representation learning framework that captures multi-domain disease profiles through patient embeddings. SCORE employs a Poisson-Adapted Latent factor Mixture (PALM) Model with pre-trained code embeddings to characterize codified features and extract meaningful patient phenotypes and embeddings. To handle the computational challenges of large-scale data, it introduces a hybrid Expectation-Maximization (EM) and Gaussian Variational Approximation (GVA) algorithm, leveraging limited labeled data to refine estimates on a vast pool of unlabeled samples. We theoretically establish the convergence of this hybrid approach, quantify GVA errors, and derive SCORE's error rate under diverging embedding dimensions. Our analysis shows that incorporating unlabeled data enhances accuracy and reduces sensitivity to label scarcity. Extensive simulations confirm SCORE's superior finite-sample performance over existing methods. Finally, we apply SCORE to predict disability status for patients with multiple sclerosis (MS) using partially labeled EHR data, demonstrating that it produces more informative and predictive patient embeddings for multiple MS-related conditions compared to existing approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。