构建首个系统化多语言个人身份信息检测基准,精准揭示模型在高敏感数据上的失效机制
REDACT: A Systematically Controlled Multilingual Benchmark for Personal Information Detection

- 通过9个控制轴生成13,427条带标注数据,覆盖51种实体类型与25种语言
- 发现规则型模型在高敏感数据上召回率仅0.07,而大模型更稳健
- 适合评估多语言隐私检测模型、研究敏感性标注对性能的影响
当前个人身份信息(PII)检测的基准设施仍显不足:现有语料库涵盖实体类型有限,生成条件随意,且无法揭示导致检测器失败的表面形式。本文提出REDACT,一个系统化控制的多语言PII基准,包含13,427条记录、324,078个实体标注、51种实体类型、4,127种表面形式模式,覆盖25种语言及9种书写系统。采用强度为2的覆盖数组采样器控制九个生成维度:领域、格式、难度、长度、密度、代码切换、语言、邻近性和共现性。三个实体级元数据字段(披露状态、披露形式、符合GDPR的敏感度层级)支持超越整体或按类型计算的F1值的分层评估。从完整基准中,我们对五种检测器(Presidio、GLiNER、OpenAI隐私过滤器、GPT-4.1和Claude Sonnet 4.6)在锁定的语言分层样本(1,000条记录)上进行评估。总体F1值掩盖了架构相关的失败结构:规则型检测器在最高风险数据上表现差,包括高敏感类别(召回率0.07)和非字面披露形式;而大模型检测器则保持更强鲁棒性,高敏感层级反而是其最强表现区间。三模型无参考大模型评判评估确认,敏感度层级划分是该任务最困难的维度。我们公开基准、数据模式、提示词及分层评估工具。
原文摘要 · Abstract (English)
Benchmark infrastructure for personally identifiable information (PII) detection remains limited: existing corpora cover few entity types, use ad hoc generation conditions, and do not show which surface conditions cause detector failures. We present REDACT, a systematically controlled multilingual PII benchmark with 13,427 records, 324,078 entity annotations, 51 entity types, 4,127 surface-form patterns, and 25 languages across 9 scripts. A strength-2 covering-array sampler controls nine generation axes: domain, format, difficulty, length, density, code-switching, language, adjacency, and co-occurrence. Three entity-level metadata fields (disclosure status, disclosure form, and a GDPR-aligned sensitivity tier) enable stratified evaluation beyond aggregate or per-type F1. From the full benchmark, we evaluate five detectors (Presidio, GLiNER, the OpenAI Privacy Filter, GPT-4.1, and Claude Sonnet 4.6) on a locked, language-stratified sample of 1,000 records. Aggregate F1 masks an architecture-dependent failure structure: the rule-based detector performs poorly on the highest-stakes data, including HIGH-sensitivity categories (recall 0.07) and non-verbatim disclosure forms, while the LLM detectors remain more robust, with the HIGH tier as their strongest sensitivity slice. A three-model reference-free LLM-as-judge assessment corroborates that sensitivity-tier assignment is the task's hardest axis. We release the benchmark, schema, prompts, and stratified evaluation harness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。