arXiv:2606.16763cs.CRcs.IT2026-06

揭示了在局部差分隐私下跨数据孤岛匿名化失效的临界条件。

Cross-Silo De-Anonymization Under Local Differential Privacy: Threat Model, Phase Transition, and Coordination Necessity

  • 构建跨孤岛个体级隐私模型,分析多源数据联合泄露风险。
  • 发现去匿名化存在相变阈值,当数据源数超临界值时攻击必然成功。
  • 证明非协同机制下跨孤岛攻击不可避免,强调协同防护必要性。

当一个人的记录分布在 k 个独立数据孤岛中,每个孤岛采用 (ε, δ)-差分隐私保护时,标准组合法则给出整体 (kε, kδ)-DP 保证。然而,该最坏情况界限无法回答实际问题:在何种条件下攻击者能真正识别目标个体?本文建立信息论框架解答此问题。提出跨孤岛个体级差分隐私(XSP-DP),其邻近关系同时捕获单个个体在所有孤岛中的记录,并验证标准基本组合律适用于该模型。在此框架下,证明去匿名化存在相变:当 k* = Θ(log n / ε²) 时(n 为总体规模,ε 为每孤岛随机响应参数),若 k << k*,Fano 下界表明任何估计器均失败;若 k >> k*,最大似然上界表明攻击成功。一个显式的异或+随机响应构造展示信息协同效应:各孤岛输出单独对目标无信息,但联合互信息严格大于零。对于非协同二元随机响应机制,证明一旦 k 超过阈值,去匿名化必然发生,确立跨孤岛协同必要性。这些结果为局部差分隐私下的跨孤岛推断攻击提供了基准威胁模型与 Θ 级阈值。

原文摘要 · Abstract (English)

When a person's records appear in k independent data silos, each protected by (epsilon, delta)-differential privacy, standard composition yields a valid (k*epsilon, k*delta)-DP guarantee for the joint output. This worst-case bound, however, does not answer the concrete inference question: at what k can an adversary actually identify a target person? This paper develops the information-theoretic framework needed to answer that question. We introduce cross-silo person-level DP (XSP-DP), a Pufferfish-style privacy notion whose adjacency relation captures all records of a single person across all silos simultaneously, and verify that the standard basic composition bound carries over to this adjacency model. Within this framework we prove that de-anonymization undergoes a phase transition at k* = Theta(log n / epsilon^2) (population size n, per-silo RR parameter epsilon): a Fano lower bound shows any estimator fails for k << k*, while a matching maximum-likelihood upper bound shows the attack succeeds for k >> k*. An explicit XOR + randomized-response construction demonstrates information synergy: each silo's output is individually uninformative about the target, yet the joint mutual information is strictly positive. For non-coordinated binary randomized-response mechanisms, we prove that de-anonymization is inevitable once k exceeds the threshold, establishing that cross-silo coordination is necessary. These results provide a baseline threat model and Theta-level threshold for cross-silo inference attacks under local DP.

差分隐私去匿名化信息论数据安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。