通过推理与排序聚合,提升文本去标识化后的身份复原攻击效果
Stronger Re-identification Attacks through Reasoning and Aggregation
- 按不同顺序识别敏感信息并融合结果,提升准确性
- 引入推理模型后,攻击成功率显著提高,尤其在背景知识丰富时
- 适合研究隐私安全、对抗攻击或数据保护的学者参考
文本去标识化技术常用于隐藏文档中的个人身份信息(PII)。然而,这些方法能否有效隐藏个体身份难以衡量。近期研究提出通过反向操作——即利用自动化攻击者基于背景知识还原被掩蔽的PII——来评估去标识化方法的鲁棒性。本文提出两种互补策略以构建更强的重识别攻击:首先发现(1)识别PII片段的顺序至关重要,对多种顺序的预测进行聚合可提升性能;其次发现(2)推理模型能显著增强攻击效果,尤其当攻击者具备丰富背景知识时。
原文摘要 · Abstract (English)
Text de-identification techniques are often used to mask personally identifiable information (PII) from documents. Their ability to conceal the identity of the individuals mentioned in a text is, however, hard to measure. Recent work has shown how the robustness of de-identification methods could be assessed by attempting the reverse process of _re-identification_, based on an automated adversary using its background knowledge to uncover the PIIs that have been masked. This paper presents two complementary strategies to build stronger re-identification attacks. We first show that (1) the _order_ in which the PII spans are re-identified matters, and that aggregating predictions across multiple orderings leads to improved results. We also find that (2) reasoning models can boost the re-identification performance, especially when the adversary is assumed to have access to extensive background knowledge.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。