提出新风险评估方法,衡量合成数据被用来推断敏感信息的可能性。
RAPID: Risk of Attribute Prediction-Induced Disclosure in Synthetic Microdata
- 基于攻击者仅用合成数据训练模型推断真实个体敏感属性的现实场景
- 对连续属性给出预测误差在容忍范围内的比例,对分类属性给出置信度提升比例
- 适用于任意生成器和算法,可指导阈值设定与不同合成方法对比
统计数据匿名化越来越多地依赖完全合成的微观数据,传统身份泄露指标对此类数据的参考价值降低。本文提出RAPID(属性预测引发泄露风险),一种直接量化在现实攻击模型下推断脆弱性的泄露风险度量。攻击者仅使用发布的合成数据训练预测模型,并将其应用于真实个体的准标识符上。对于连续敏感属性,RAPID报告预测值落在指定相对误差容差内的记录比例;对于分类属性,提出基线归一化的置信度得分,衡量攻击者对真实类别的信心相对于类别先验概率的提升程度,并将风险总结为超过政策定义阈值的记录比例。该方法生成可解释、有界的风险度量,对类别不平衡鲁棒,不依赖特定合成器,且可适配任意学习算法。通过模拟和真实数据展示阈值校准、不确定性量化及合成数据生成器的比较评估。结果表明,RAPID提供了属性推断泄露风险的实际、贴近攻击者的上限估计,补充了现有效用诊断与披露控制框架。
原文摘要 · Abstract (English)
Statistical data anonymization increasingly relies on fully synthetic microdata, for which classical identity disclosure measures are less informative than an adversary's ability to infer sensitive attributes from released data. We introduce RAPID (Risk of Attribute Prediction--Induced Disclosure), a disclosure risk measure that directly quantifies inferential vulnerability under a realistic attack model. An adversary trains a predictive model solely on the released synthetic data and applies it to real individuals' quasi-identifiers. For continuous sensitive attributes, RAPID reports the proportion of records whose predicted values fall within a specified relative error tolerance. For categorical attributes, we propose a baseline-normalized confidence score that measures how much more confident the attacker is about the true class than would be expected from class prevalence alone, and we summarize risk as the fraction of records exceeding a policy-defined threshold. This construction yields an interpretable, bounded risk metric that is robust to class imbalance, independent of any specific synthesizer, and applicable with arbitrary learning algorithms. We illustrate threshold calibration, uncertainty quantification, and comparative evaluation of synthetic data generators using simulations and real data. Our results show that RAPID provides a practical, attacker-realistic upper bound on attribute-inference disclosure risk that complements existing utility diagnostics and disclosure control frameworks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。