质疑现有隐私信息移除攻击评估的可信度,指出数据泄露导致结果虚高。
The Conundrum of Trustworthy Research on Attacking Personally Identifiable Information Removal Techniques
- 通过分析攻击评估方法,发现数据泄露和污染问题普遍存在。
- 真实隐私数据难获取,现有实验无法反映实际防护效果。
- 呼吁建立无数据泄露的评测标准,适合隐私安全研究者参考。
从文本中移除个人身份信息(PII)是遵守数据保护法规、实现数据共享而不泄露隐私的必要手段。然而,近期研究显示经PII移除处理的文档仍易受重建攻击。我们怀疑这些攻击的成功率被严重夸大。通过对现有攻击评估的批判性分析,发现数据泄露和数据污染未被妥善处理,导致无法客观判断PII移除技术在真实场景下的隐私保护能力。我们探讨了可能的无泄露数据源与攻击设置,结论是只有真正私有的数据才能进行客观评估。但出于合理原因,私人数据访问受到严格限制,这使得公开研究社区难以以透明、可复现、可信的方式解决该问题。
原文摘要 · Abstract (English)
Removing personally identifiable information (PII) from texts is necessary to comply with various data protection regulations and to enable data sharing without compromising privacy. However, recent works show that documents sanitized by PII removal techniques are vulnerable to reconstruction attacks. Yet, we suspect that the reported success of these attacks is largely overestimated. We critically analyze the evaluation of existing attacks and find that data leakage and data contamination are not properly mitigated, leaving the question whether or not PII removal techniques truly protect privacy in real-world scenarios unaddressed. We investigate possible data sources and attack setups that avoid data leakage and conclude that only truly private data can allow us to objectively evaluate vulnerabilities in PII removal techniques. However, access to private data is heavily restricted - and for good reasons - which also means that the public research community cannot address this problem in a transparent, reproducible, and trustworthy manner.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。