用唯一性比率提前预警数据合并后的隐私泄露风险
Uniqueness ratio as a predictor of a privacy leakage
- 通过计算合并前属性组合的唯一性比率预判风险
- 高唯一性比率与合并后记录可唯一识别比例显著相关
- 为数据工程师提供简单可解释的隐私风险预警工具
当独立数据库在各自匿名化后合并时,仍可能引发身份泄露。现有研究多关注合并后的检测或复杂隐私模型,而忽视了合并前简单可解释的预警指标。本研究探索候选连接属性的唯一性比率作为再识别风险的早期预测因子。基于合成多表数据集,计算各数据库中属性组合的唯一性比率,并分析其与合并后身份暴露程度的相关性。实验结果表明,合并前的高唯一性比率与合并后记录可唯一识别或落入极小群体的比例显著正相关。研究证实,唯一性比率可提供可解释且实用的隐私风险信号,为构建更全面的合并前风险评估模型奠定基础。
原文摘要 · Abstract (English)
Identity leakage can emerge when independent databases are joined, even when each dataset is anonymized individually. While previous work focuses on post-join detection or complex privacy models, little attention has been given to simple, interpretable pre-join indicators that can warn data engineers and database administrators before integration occurs. This study investigates the uniqueness ratio of candidate join attributes as an early predictor of re-identification risk. Using synthetic multi-table datasets, we compute the uniqueness ratio of attribute combinations within each database and examine how these ratios correlate with identity exposure after the join. Experimental results show a strong relationship between high pre-join uniqueness and increased post-join leakage, measured by the proportion of records that become uniquely identifiable or fall into very small groups. Our findings demonstrate that uniqueness ratio offers an explainable and practical signal for assessing join induced privacy risk, providing a foundation for developing more comprehensive pre-join risk estimation models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。