构建首个关注上下文隐私的敏感信息擦除评测集,揭示当前模型在真实场景中的严重不足。
RedactionBench

- 基于上下文完整性理论,设计涵盖11个领域的200份真实文档评测集
- 提出字符级R-Score指标,消除格式干扰并公平评估语义相似擦除
- 发现模型对上下文相关擦除共识仅47.7%,适合隐私安全与AI伦理研究者
大型语言模型越来越多地应用于需擦除个人身份信息(PII)的敏感领域。现有评测集混淆了信息提取机制与隐私语义。公共电话号码与医疗记录中的电话号码具有本质区别,是否构成泄露取决于持有者、目的及上下文,这使擦除远超简单实体识别。基于上下文完整性理论,我们构建了红化基准测试(RedactionBench),包含200份跨11个领域的多样化文档,主要源自真实数据。同时提出R-Score,一种新型字符级指标,同等对待语义相似的擦除,并消除如掩码样式等表面格式差异的影响。在命名实体识别模型、小规模语言模型及配备代理工具的前沿模型上进行评估,结果表明上下文擦除仍是未解难题。对超过80名用户的真人评估显示,对强制擦除(89.4%)和安全文本保留(94.1%)有高度共识,但对上下文擦除仅47.7%一致。该差异凸显上下文隐私的主观性,推动了R-Score的设计,使其将上下文模糊性与严格精度分离。我们对比35个模型家族的表现,并公开发布RedactionBench,以建立未来隐私保护系统评估基线,期望激发高效模型设计与标准化评测。
原文摘要 · Abstract (English)
Large Language Models are increasingly applied to sensitive domains that require redaction of personally identifiable information (PII). While redacting PII is a data cleaning prerequisite, existing benchmarks conflate extraction mechanics with privacy semantics. A public phone number is not equivalent to a phone number in a medical record. Whether information constitutes a violation depends heavily on who holds it, why, and in what context, fundamentally differentiating redaction from simple entity recognition. Grounded in contextual integrity, we introduce RedactionBench, a manually annotated benchmark comprising 200 diverse documents across 11 domains, mostly seeded from real-world sources. We also introduce R-Score, a novel character-level metric that treats semantically similar redactions equally and nullifies shallow formatting choices, such as varying masking styles for phone numbers. Evaluations across Named Entity Recognition models, entity extraction Small Language Models, and frontier models equipped with agentic tools demonstrate that contextual redaction remains an unsolved problem. A human evaluation with over 80 users on RedactionBench reveals a stark dichotomy in privacy perceptions. Annotators show consensus with target labels for mandatory redactions (89.4 percent) and safe text preservations (94.1 percent), but fail to agree on contextual redactions (47.7 percent). This variance demonstrates the subjective nature of contextual privacy and motivates R-Score, which decouples contextual ambiguity from strict precision. We compare 35 models across families and report their performance in redacting PII. Finally, we release RedactionBench to establish a baseline for future privacy-preserving systems, hoping to inspire efficient model design and standardized evaluations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。