基于联邦学习生成时序数据,保护隐私且保持数据可用性
VFLGAN-TS: Vertical Federated Learning-based Generative Adversarial Networks for Publication of Vertically Partitioned Time-Series Data
- 用属性判别器+垂直联邦学习生成跨方时序数据
- 生成效果接近集中式训练上限,误差<5%
- 支持差分隐私并内置隐私审计机制
在当前人工智能时代,数据规模与质量对训练高质量模型至关重要。然而,由于隐私顾虑和法规限制,原始数据往往无法共享。一种潜在解决方案是发布与私有数据分布相似的合成数据集。但在某些场景中,训练所需属性分布在不同参与方之间,各方因隐私法规无法共享本地数据以构建合成数据。此前我们提出首个用于静态数据的垂直联邦生成对抗网络(VFLGAN),但该方法难以有效处理同时具备时间维度和属性维度的时序数据。本文提出VFLGAN-TS,结合属性判别器与垂直联邦学习,在垂直分区场景下生成合成时序数据。其性能接近集中式训练的上限,差异小于5%。为进一步保护隐私,我们引入高斯机制,使VFLGAN-TS满足(ε,δ)-差分隐私。此外,我们设计了一种增强型隐私审计方案,评估框架内合成数据可能引发的隐私泄露风险。
原文摘要 · Abstract (English)
In the current artificial intelligence (AI) era, the scale and quality of the dataset play a crucial role in training a high-quality AI model. However, often original data cannot be shared due to privacy concerns and regulations. A potential solution is to release a synthetic dataset with a similar distribution to the private dataset. Nevertheless, in some scenarios, the attributes required to train an AI model are distributed among different parties, and the parties cannot share the local data for synthetic data construction due to privacy regulations. In PETS 2024, we recently introduced the first Vertical Federated Learning-based Generative Adversarial Network (VFLGAN) for publishing vertically partitioned static data. However, VFLGAN cannot effectively handle time-series data, presenting both temporal and attribute dimensions. In this article, we proposed VFLGAN-TS, which combines the ideas of attribute discriminator and vertical federated learning to generate synthetic time-series data in the vertically partitioned scenario. The performance of VFLGAN-TS is close to that of its counterpart, which is trained in a centralized manner and represents the upper limit for VFLGAN-TS. To further protect privacy, we apply a Gaussian mechanism to make VFLGAN-TS satisfy an $(ε,δ)$-differential privacy. Besides, we develop an enhanced privacy auditing scheme to evaluate the potential privacy breach through the framework of VFLGAN-TS and synthetic datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。