联邦学习+预测推断,私有数据也能做可靠统计推断
Federated Prediction-Powered Inference from Decentralized Data
- 用联邦学习聚合本地模型,不共享原始数据
- 在私有数据上仍能生成有效的置信区间
- 适合医疗、金融等隐私敏感领域使用
在多个领域,机器学习的广泛应用使得研究人员能够获取低成本的预测数据,可作为统计推断的辅助数据。尽管这些数据相比金标准数据集可靠性较低,但预测推断(PPI)已被提出以确保在不可靠数据下仍保持统计有效性。然而,当私有的金标准数据因隐私无法共享用于模型训练时,便产生‘数据孤岛’问题,导致预测模型准确性下降,推断结果失效。本文提出联邦预测推断(Fed-PPI)框架,通过在分布式数据上训练本地模型,利用联邦学习(FL)聚合模型,并结合PPI计算方法,实现无需共享私有信息即可得到统计有效的结论。实验验证了该框架在生成有效置信区间方面的有效性。
原文摘要 · Abstract (English)
In various domains, the increasing application of machine learning allows researchers to access inexpensive predictive data, which can be utilized as auxiliary data for statistical inference. Although such data are often unreliable compared to gold-standard datasets, Prediction-Powered Inference (PPI) has been proposed to ensure statistical validity despite the unreliability. However, the challenge of `data silos' arises when the private gold-standard datasets are non-shareable for model training, leading to less accurate predictive models and invalid inferences. In this paper, we introduces the Federated Prediction-Powered Inference (Fed-PPI) framework, which addresses this challenge by enabling decentralized experimental data to contribute to statistically valid conclusions without sharing private information. The Fed-PPI framework involves training local models on private data, aggregating them through Federated Learning (FL), and deriving confidence intervals using PPI computation. The proposed framework is evaluated through experiments, demonstrating its effectiveness in producing valid confidence intervals.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。