arXiv:2511.02754stat.MEcs.LG2025-11

用分布式方法提升医疗数据的高效低隐私风险表示学习

DANIEL: A Distributed and Scalable Approach for Global Representation Learning with EHR Applications

  • 基于伊辛模型设计分布式优化框架,支持大规模二值数据建模
  • 在58,248例患者数据上实现更优的临床表型与关系识别效果
  • 适合需要跨机构协作且保护数据隐私的研究者使用

传统概率图模型在高维、异构且受数据共享限制的现代数据环境中面临根本性挑战。本文重新审视马尔可夫随机场家族中的经典伊辛模型,提出一种分布式框架,可在保留低秩结构的前提下,实现大规模二值数据的可扩展、隐私保护表示学习。该方法通过双因子梯度下降优化非凸代理损失函数,在计算与通信效率上显著优于传统凸方法。我们在匹兹堡大学医学中心(UPMC)与麻省总医院(MGB)的多机构电子健康记录(EHR)数据集上进行评估,覆盖58,248名患者,结果表明该算法在全局表示学习及下游临床任务(包括关系检测、患者表型分类和聚类)中表现优异。这些成果展示了联邦高维场景下统计推断的广泛潜力,同时应对了数据复杂性与多机构整合的实际挑战。

原文摘要 · Abstract (English)

Classical probabilistic graphical models face fundamental challenges in modern data environments, which are characterized by high dimensionality, source heterogeneity, and stringent data-sharing constraints. In this work, we revisit the Ising model, a well-established member of the Markov Random Field (MRF) family, and develop a distributed framework that enables scalable and privacy-preserving representation learning from large-scale binary data with inherent low-rank structure. Our approach optimizes a non-convex surrogate loss function via bi-factored gradient descent, offering substantial computational and communication advantages over conventional convex approaches. We evaluate our algorithm on multi-institutional electronic health record (EHR) datasets from 58,248 patients across the University of Pittsburgh Medical Center (UPMC) and Mass General Brigham (MGB), demonstrating superior performance in global representation learning and downstream clinical tasks, including relationship detection, patient phenotyping, and patient clustering. These results highlight a broader potential for statistical inference in federated, high-dimensional settings while addressing the practical challenges of data complexity and multi-institutional integration.

医疗人工智能联邦学习表示学习低秩建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。