arXiv:2606.08372cs.CRcs.LG2026-06

系统化分析合成表格数据的重构攻击,揭示隐私风险关键因素。

SoK: Reconstruction Attacks on Synthetic Tabular Data (Insights from Winning the NIST CRC)

论文配图:SoK: Reconstruction Attacks on Synthetic Tabular Data (Insights from Winning the NIST CRC)
图 1 · 摘自论文原文
  • 构建攻击分类体系,覆盖14种攻击与9种生成方法
  • 发现差分隐私仅在ε≤1时有效,且受生成器能力限制
  • 多数攻击源于分布结构而非记忆,异常记录风险更高

合成数据被视为保护敏感表格记录的隐私替代方案,但其核心威胁——重构攻击(即从合成数据和少量已知准标识符中恢复个体隐藏属性值)此前研究零散且难以比较。本文首次系统化梳理了针对去标识化与合成表格数据的重构(等价于属性推断)攻击。我们提出一种基于攻击所利用结构的分类体系;开展迄今最系统的实证评估,将14种攻击对阵9种合成数据生成(SDG)方法,在5个基准数据集上测试;并提出新攻击填补分类空白,其中一种(CoBP-RA)为最强攻击。我们引入解读攻击成功的新方法:通过记忆测试区分对总体分布的重构与训练记录的记忆;并通过归约将重构与成员推断置于可比较尺度。结果表明:选择何种SDG方法比选择何种攻击影响更大;差分隐私仅在小预算(ε≲1)下提供保护,之后保护效果趋于饱和,受生成器容量限制而非噪声水平;去标识化方法最为脆弱;大多数重构反映的是分布结构而非记忆,个体风险集中于异常记录。所提攻击与基础设施经2025年美国国家标准与技术研究院(NIST)协同研究周期中所有红队竞赛第一的成绩验证。

原文摘要 · Abstract (English)

Synthetic data is increasingly promoted as a privacy-preserving substitute for releasing sensitive tabular records, yet its central adversarial threat ("reconstruction", the recovery of an individual's hidden attribute values from a synthetic release and a handful of known quasi-identifiers) has been studied only in scattered, hard-to-compare settings. We present the first systematization of reconstruction (equivalently, attribute inference) attacks on de-identified and synthetic tabular data. We contribute a taxonomy that organizes attacks by the structure they exploit; the most systematic empirical evaluation to date, pitting fourteen attacks against nine synthetic data generation (SDG) methods across five benchmark datasets; and a set of new attacks that fill gaps in the taxonomy, one of which (CoBP-RA) is the strongest attack we measure. Crucially, we introduce a methodology for interpreting what attack success means: a memorization test that distinguishes reconstruction of the population distribution from memorization of training records, and a reduction that places reconstruction and membership inference on a single comparable scale. Our findings: the choice of SDG method governs risk far more than the choice of attack; differential privacy protects mainly at small budgets ($\varepsilon\lesssim1$), above which protection plateaus, bounded by the synthesizer's capacity rather than its noise; de-identification methods are the most exposed; and most reconstruction reflects distributional structure rather than memorization, concentrating individual risk on atypical records. The attacks and infrastructure are externally validated by our first-place finish among all red teams in the 2025 \textit{National Institute of Standards and Technology} (NIST) Collaborative Research Cycle.

合成数据隐私攻击差分隐私数据安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。