联邦聚类中公开中心点会暴露原始数据,攻击者可完美重建输入。
Breaking Privacy in Federated Clustering: Perfect Input Reconstruction via Temporal Correlations
- 利用聚类迭代中的时间规律,结合分配信息进行代数分析
- 在真实场景下实现对原始数据的完全重构,无需额外假设
- 揭示效率与隐私的根本矛盾,适合关注安全性的研究者
联邦聚类允许多方在不共享原始数据的情况下发现分布式数据中的模式。为降低开销,许多协议在训练过程中披露中间聚类中心。尽管常被视为高效但无害,这种披露是否威胁隐私仍存疑。先前分析将问题建模为隐藏子集和问题(HSSP),认为经典攻击无法恢复输入,故认为安全。我们重新审视该问题,发现k-means迭代中的时间规律会产生可被利用的结构,从而实现完全输入重建。基于此,提出轨迹感知重建(TAR)攻击,结合时间分配信息与代数分析,精确还原原始数据。结果首次以实际攻击证明:联邦聚类中披露聚类中心会显著损害隐私,暴露效率与隐私之间的根本矛盾。
原文摘要 · Abstract (English)
Federated clustering allows multiple parties to discover patterns in distributed data without sharing raw samples. To reduce overhead, many protocols disclose intermediate centroids during training. While often treated as harmless for efficiency, whether such disclosure compromises privacy remains an open question. Prior analyses modeled the problem as a so-called Hidden Subset Sum Problem (HSSP) and argued that centroid release may be safe, since classical HSSP attacks fail to recover inputs. We revisit this question and uncover a new leakage mechanism: temporal regularities in $k$-means iterations create exploitable structure that enables perfect input reconstruction. Building on this insight, we propose Trajectory-Aware Reconstruction (TAR), an attack that combines temporal assignment information with algebraic analysis to recover exact original inputs. Our findings provide the first rigorous evidence, supported by a practical attack, that centroid disclosure in federated clustering significantly compromises privacy, exposing a fundamental tension between privacy and efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。