通过噪声锚点对齐,实现高效隐私保护的协作学习。
Geometric Data Perturbation with Noisy-Anchor Alignment for Privacy-Preserving Collaborative Learning

- 将噪声注入共享锚点而非数据,提升隐私性
- 在MNIST和CelebA上实现更高学习准确率
- 适合关注隐私与性能平衡的联邦学习研究者
几何数据扰动(GDP)支持一次性隐私保护协作学习:各参与者对其私有数据施加保持距离的变换后,仅上传结果表示给中心分析师。研究分析了分析师与参与者共谋情形下,攻击者利用上传表示与共谋者披露的变换和数据恢复非共谋者私有数据的风险。独立变换虽抵抗攻击,但导致表示空间不兼容,降低下游模型性能。共享锚点对齐可恢复兼容性并提升效用,但披露锚矩阵会暴露非共谋者数据。直接对数据表示加噪虽缓解漏洞,但严重损害效用。本文提出向锚表示加噪:参与者独立变换私有数据与共享锚,仅扰动锚表示并单轮上传。分析师使用带噪锚表示通过广义正交投影问题对齐数据表示。理论分析对齐与恢复误差,并给出收敛的保守充分条件;评估三类恢复攻击。在MNIST和CelebA上的实验表明,相比数据加噪,在相似泄漏水平下,锚噪声实现更高学习准确率,提供更优隐私-效用权衡。
原文摘要 · Abstract (English)
Geometric Data Perturbation (GDP) enables one-shot, privacy-preserving collaborative learning: each participant applies a distance-preserving transformation to its private data and uploads only the resulting representation to a central analyst. We study GDP under analyst-participant collusion, in which the analyst combines all uploaded representations with the private data and transformations disclosed by colluding participants to recover a non-colluding participant's private data. Participant-specific independent transformations resist this attack but map participants' data into incompatible representation spaces, degrading downstream model performance. Shared-anchor alignment from Data Collaboration (DC) analysis restores compatibility and improves utility, but we show that disclosing the DC anchor matrix enables exact recovery of non-colluding participants' private data even in the presence of collusion. Adding noise directly to the private-data representations mitigates this vulnerability but substantially reduces utility. We propose adding noise to the anchor representations instead. Each participant independently transforms its private data and the shared anchor matrix, perturbs only the resulting anchor representation, and uploads both representations in a single round. Using the noisy anchor representations, the analyst aligns the private-data representations by solving a Generalized Orthogonal Procrustes Problem. We characterize alignment and recovery errors, specialize a conservative sufficient condition for convergence of the alignment to our setting, and analyze three recovery attacks. Experiments on MNIST and CelebA show that, across the evaluated attacks and deployment settings, anchor noise achieves higher learning accuracy than private-data noise at comparable measured leakage, yielding a more favorable privacy-utility trade-off under the specified collusion model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。