arXiv:2411.17287cs.LG2024-11

用高斯过程实现生物数据回归的隐私保护联邦域适应。

Privacy-Preserving Federated Unsupervised Domain Adaptation for Regression on Small-Scale and High-Dimensional Biological Data

  • 基于高斯过程与随机编码,实现无监督回归的联邦训练。
  • 在DNA甲基化年龄预测任务中达到中心化最优水平。
  • 适合小规模、高维度、跨机构的生物数据隐私建模。

机器学习模型在小规模异构数据集上常因数据采集差异和人群差异导致的领域偏移而难以泛化。这一问题在生物数据中尤为突出,因其具有高维度、小样本、分布式存储的特点。现有联邦域适应方法多依赖深度学习且集中于分类任务,不适用于此类场景。本文提出freda,一种用于回归任务的隐私保护联邦无监督域适应方法。不同于基于深度学习的FDA方法,freda首次实现通过随机编码与安全聚合,在不直接访问原始数据的前提下,联邦训练高斯过程以建模复杂特征关系。该方法在挑战性的DNA甲基化数据年龄预测任务中表现优异,性能媲美集中式最先进方法,同时保障完全数据隐私。

原文摘要 · Abstract (English)

Machine learning models often struggle with generalization in small, heterogeneous datasets due to domain shifts caused by variations in data collection and population differences. This challenge is particularly pronounced in biological data, where data is high-dimensional, small-scale, and decentralized across institutions. While federated domain adaptation methods (FDA) aim to address these challenges, most existing approaches rely on deep learning and focus on classification tasks, making them unsuitable for small-scale, high-dimensional applications. In this work, we propose freda, a privacy-preserving federated method for unsupervised domain adaptation in regression tasks. Unlike deep learning-based FDA approaches, freda is the first method to enable the federated training of Gaussian Processes to model complex feature relationships while ensuring complete data privacy through randomized encoding and secure aggregation. This allows for effective domain adaptation without direct access to raw data, making it well-suited for applications involving high-dimensional, heterogeneous datasets. We evaluate freda on the challenging task of age prediction from DNA methylation data, demonstrating that it achieves performance comparable to the centralized state-of-the-art method while preserving complete data privacy.

联邦学习高斯过程生物信息隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。