arXiv:2602.07126cs.LG2026-02被引 1

提出新型隐私攻击,检测多表合成数据中用户实体的泄露风险。

Finding Connections: Membership Inference Attacks for the Multi-Table Synthetic Data Setting

  • 基于异构图神经网络,从多表关联信息中构建用户实体表示。
  • 在真实数据集上验证,顶尖生成模型仍存在用户级隐私泄露。
  • 适用于评估合成数据生成器的用户隐私保护能力,尤其关注关系型数据。

合成表格数据因其支持隐私保护的数据共享而受到关注。尽管在单表合成生成方面已取得进展(以行或项目为建模单位),但多数现实世界数据存在于关系型数据库中,用户的个人信息跨越多个相互连接的表。近期合成关系型数据生成技术应运而生以应对这一复杂性,但其发布引入了独特的隐私挑战:信息不仅可能从单个项目泄露,还可能通过构成完整用户实体的表间关系泄露。为此,我们提出一种新的会员推理攻击(MIA)设置,用于审计合成关系型数据的用户级隐私,并表明现有的单表MIA在项目层面审计会低估用户级隐私泄露。我们进一步提出多表会员推理攻击(MT-MIA),这是一种在无箱威胁模型下的新型对抗攻击,通过异构图神经网络针对学习到的用户实体表示进行攻击。通过整合一个用户的所有相关项目,MT-MIA比现有攻击更有效地针对由表间关系引发的用户级漏洞。我们在多种真实世界多表数据集上评估了MT-MIA,证明这种漏洞存在于最先进的关系型合成数据生成器中,并利用该攻击进一步分析泄露发生的位置。

原文摘要 · Abstract (English)

Synthetic tabular data has gained attention for enabling privacy-preserving data sharing. While substantial progress has been made in single-table synthetic generation where data are modeled at the row or item level, most real-world data exists in relational databases where a user's information spans items across multiple interconnected tables. Recent advances in synthetic relational data generation have emerged to address this complexity, yet release of these data introduce unique privacy challenges as information can be leaked not only from individual items but also through the relationships that comprise a complete user entity. To address this, we propose a novel Membership Inference Attack (MIA) setting to audit the empirical user-level privacy of synthetic relational data and show that single-table MIAs that audit at an item level underestimate user-level privacy leakage. We then propose Multi-Table Membership Inference Attack (MT-MIA), a novel adversarial attack under a No-Box threat model that targets learned representations of user entities via Heterogeneous Graph Neural Networks. By incorporating all connected items for a user, MT-MIA better targets user-level vulnerabilities induced by inter-tabular relationships than existing attacks. We evaluate MT-MIA on a range of real-world multi-table datasets and demonstrate that this vulnerability exists in state-of-the-art relational synthetic data generators, employing MT-MIA to additionally study where this leakage occurs.

隐私攻击合成数据关系型数据图神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。