arXiv:2504.21199stat.MLcs.CR2025-04

从有限统计数据中推断出必然正确的数据子集,揭示隐私泄露风险。

Generate-then-Verify: Reconstructing Data from Limited Published Statistics

  • 先生成候选数据结论,再验证其在所有可能数据集中是否恒成立
  • 在普查住房数据上验证,即使数据稀疏仍存在可确证的隐私泄露
  • 适用于研究统计发布中的隐私脆弱性,尤其关注数据重构边界

我们研究从汇总统计中重构表格数据的问题,攻击者的目标是识别出可被汇总数据100%验证的敏感数据结论。以往工作在统计量丰富时可完全重建数据,但本文聚焦于统计量不足、多个数据集均符合公布统计的情形,此时无法完美重建整个数据集。我们提出部分数据重构问题:攻击者只需输出一个保证正确的行或列子集。为此,我们设计一种新的整数规划方法,先生成候选结论,再验证其在所有与已公布统计一致的数据集中是否恒成立。我们在美国十年一次人口普查的住房级微观数据上评估该方法,发现即便发布的数据相对稀疏,隐私泄露仍可能持续存在。

原文摘要 · Abstract (English)

We study the problem of reconstructing tabular data from aggregate statistics, in which the attacker aims to identify interesting claims about the sensitive data that can be verified with 100% certainty given the aggregates. Successful attempts in prior work have conducted studies in settings where the set of published statistics is rich enough that entire datasets can be reconstructed with certainty. In our work, we instead focus on the regime where many possible datasets match the published statistics, making it impossible to reconstruct the entire private dataset perfectly (i.e., when approaches in prior work fail). We propose the problem of partial data reconstruction, in which the goal of the adversary is to instead output a $\textit{subset}$ of rows and/or columns that are $\textit{guaranteed to be correct}$. We introduce a novel integer programming approach that first $\textbf{generates}$ a set of claims and then $\textbf{verifies}$ whether each claim holds for all possible datasets consistent with the published aggregates. We evaluate our approach on the housing-level microdata from the U.S. Decennial Census release, demonstrating that privacy violations can still persist even when information published about such data is relatively sparse.

数据重构隐私保护统计披露

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。