arXiv:2510.09184cs.CL2025-10

通过推理与排序聚合,提升文本去标识化后的身份复原攻击效果

Stronger Re-identification Attacks through Reasoning and Aggregation

  • 按不同顺序识别敏感信息并融合结果,提升准确性
  • 引入推理模型后,攻击成功率显著提高,尤其在背景知识丰富时
  • 适合研究隐私安全、对抗攻击或数据保护的学者参考

文本去标识化技术常用于隐藏文档中的个人身份信息(PII)。然而,这些方法能否有效隐藏个体身份难以衡量。近期研究提出通过反向操作——即利用自动化攻击者基于背景知识还原被掩蔽的PII——来评估去标识化方法的鲁棒性。本文提出两种互补策略以构建更强的重识别攻击:首先发现(1)识别PII片段的顺序至关重要,对多种顺序的预测进行聚合可提升性能;其次发现(2)推理模型能显著增强攻击效果,尤其当攻击者具备丰富背景知识时。

原文摘要 · Abstract (English)

Text de-identification techniques are often used to mask personally identifiable information (PII) from documents. Their ability to conceal the identity of the individuals mentioned in a text is, however, hard to measure. Recent work has shown how the robustness of de-identification methods could be assessed by attempting the reverse process of _re-identification_, based on an automated adversary using its background knowledge to uncover the PIIs that have been masked. This paper presents two complementary strategies to build stronger re-identification attacks. We first show that (1) the _order_ in which the PII spans are re-identified matters, and that aggregating predictions across multiple orderings leads to improved results. We also find that (2) reasoning models can boost the re-identification performance, especially when the adversary is assumed to have access to extensive background knowledge.

身份识别隐私安全对抗攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。