arXiv:2503.12008cs.LGcs.AI2025-03被引 13

针对表格数据扩散模型,提出高效会员推理攻击方法并夺冠。

Winning the MIDST Challenge: New Membership Inference Attacks on Diffusion Models for Tabular Data Synthesis

  • 利用损失特征在不同噪声和时间步上的变化,设计轻量MLP攻击
  • 在MIDST挑战赛中所有赛道排名第一,验证了攻击有效性
  • 揭示噪声初始化对攻击效果的关键影响,适合隐私安全研究者

使用扩散模型进行表格数据合成近年来受到广泛关注,因其有望在数据效用与隐私保护间取得平衡。然而,现有隐私评估常依赖启发式指标或较弱的会员推理攻击(MIA),难以准确衡量隐私风险。本文对基于扩散模型的表格数据合成开展严谨的MIA研究,发现针对图像模型设计的先进攻击在此场景中失效。我们识别出噪声初始化是影响攻击效果的关键因素,并提出一种基于机器学习的方法,通过捕捉不同噪声和时间步下的损失特征来有效学习会员信号。该方法采用轻量级MLP实现,无需手动调参。在SaTML 2025举办的MIDST挑战赛中,本方法在所有赛道均获得第一名。代码已开源:https://github.com/Nicholas0228/Tartan_Federer_MIDST。

原文摘要 · Abstract (English)

Tabular data synthesis using diffusion models has gained significant attention for its potential to balance data utility and privacy. However, existing privacy evaluations often rely on heuristic metrics or weak membership inference attacks (MIA), leaving privacy risks inadequately assessed. In this work, we conduct a rigorous MIA study on diffusion-based tabular synthesis, revealing that state-of-the-art attacks designed for image models fail in this setting. We identify noise initialization as a key factor influencing attack efficacy and propose a machine-learning-driven approach that leverages loss features across different noises and time steps. Our method, implemented with a lightweight MLP, effectively learns membership signals, eliminating the need for manual optimization. Experimental results from the MIDST Challenge @ SaTML 2025 demonstrate the effectiveness of our approach, securing first place across all tracks. Code is available at https://github.com/Nicholas0228/Tartan_Federer_MIDST.

会员推理扩散模型表格数据隐私安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。