用知识迁移提升罕见病数据的低维表示效果
Enhancing Spectral Embedding through Robust and Flexible Knowledge Transfer in Electronic Health Records

- 通过柔性共享机制融合大人群知识,突破传统一一对应限制
- 在多发性硬化真实数据中,弱共享信号场景下性能超越现有方法
- 适合罕见病研究、小样本医疗表征学习场景
针对罕见病队列中电子健康记录数据高维但样本量有限的问题,我们提出一种基于谱的无监督表征学习框架,用于生成临床概念和患者的低维嵌入。为克服数据局限,方法引入来自更广泛人群的知识矩阵,该矩阵与罕见病队列存在部分重叠的潜在子空间。不同于现有方法强加的严格一一对应信号对齐假设,本方法允许更灵活、更真实的结构化知识共享。提出两步谱嵌入流程:首先识别并移除知识矩阵中的无关成分;随后采用投影方法分别恢复共享与异质成分。模拟实验及对真实多发性硬化队列的分析表明,该方法在共享信号微弱且仅部分对齐的挑战性场景下显著优于现有方法。
原文摘要 · Abstract (English)
We propose a spectral-based, unsupervised representation learning framework to derive low-dimensional embeddings for clinical concepts and patients in rare disease cohorts from electronic health records, where data are high-dimensional but sample sizes are limited. To overcome this challenge, we incorporate a knowledge matrix extracted from a broader population that shares a partially overlapping subspace with the rare-disease cohort. Our method departs from existing approaches by relaxing restrictive one-to-one signal-alignment assumptions between the latent data matrix and knowledge matrix, allowing more flexible and realistic forms of structured sharing. We introduce a novel two-step spectral embedding procedure: first, we identify and remove irrelevant components from the knowledge matrix; then, we apply a projection-based method to separately recover shared and heterogeneous components. Simulations and an analysis of a real-world multiple sclerosis cohort show that the proposed method outperforms competing approaches, particularly in challenging scenarios where shared signals are weak and only partially aligned, as is common in rare-disease data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。