测试主流隐私掩码模型发现严重漏检,暴露真实隐私风险。
Unmasking the Reality of PII Masking Models: Performance Gaps and the Call for Accountability
- 构建1.7万条跨域半合成数据,覆盖5类检测难点
- 模型在噪声、新实体等场景下漏检率超40%
- 呼吁建立模型卡中的上下文披露机制
隐私掩码是数据隐私中的关键概念,涉及个人身份信息(PII)的匿名化与去匿名化。该技术依赖自然语言处理中的命名实体识别(NER)方法来识别和分类文本中的命名实体。然而,现有方法存在诸多局限:内容敏感性(如歧义、多义、上下文依赖或领域特定内容)、表达变体(如昵称、别名、非正式表达、替代形式、新兴词汇及演变命名规范)以及格式或语法差异、拼写错误等。尽管已有多个广泛使用的PII数据集被用于训练模型(如Piiranha和Starpii,分别在HuggingFace上下载超30万次和58万次),但这些数据集难以应对实际复杂场景。本文收集了来自印度、英国和美国等多个司法管辖区的17,000条独特半合成句子,涵盖16类PII,并基于语言模型生成包含五类核心检测维度的语句:(1)基础实体识别,(2)上下文实体消歧,(3)噪声与真实世界数据中的NER,(4)新兴与新型实体检测,(5)跨语言或多语言NER,另含一种对抗性场景。实验结果揭示了当前模型使用所引发的隐私泄露风险(结合其累计下载量评估)。研究强调需改进模型性能评估体系,并在模型卡中增加上下文披露要求。
原文摘要 · Abstract (English)
Privacy Masking is a critical concept under data privacy involving anonymization and de-anonymization of personally identifiable information (PII). Privacy masking techniques rely on Named Entity Recognition (NER) approaches under NLP support in identifying and classifying named entities in each text. NER approaches, however, have several limitations including (a) content sensitivity including ambiguous, polysemic, context dependent or domain specific content, (b) phrasing variabilities including nicknames and alias, informal expressions, alternative representations, emerging expressions, evolving naming conventions and (c) formats or syntax variations, typos, misspellings. However, there are a couple of PII datasets that have been widely used by researchers and the open-source community to train models on PII detection or masking. These datasets have been used to train models including Piiranha and Starpii, which have been downloaded over 300k and 580k times on HuggingFace. We examine the quality of the PII masking by these models given the limitations of the datasets and of the NER approaches. We curate a dataset of 17K unique, semi-synthetic sentences containing 16 types of PII by compiling information from across multiple jurisdictions including India, U.K and U.S. We generate sentences (using language models) containing these PII at five different NER detection feature dimensions - (1) Basic Entity Recognition, (2) Contextual Entity Disambiguation, (3) NER in Noisy & Real-World Data, (4) Evolving & Novel Entities Detection and (5) Cross-Lingual or multi-lingual NER) and 1 in adversarial context. We present the results and exhibit the privacy exposure caused by such model use (considering the extent of lifetime downloads of these models). We conclude by highlighting the gaps in measuring performance of the models and the need for contextual disclosure in model cards for such models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。