arXiv:2605.09973cs.CLcs.AI2026-05被引 1

小模型精准识别42类敏感信息,支持多语言隐私保护

GLiNER2-PII: A Multilingual Model for Personally Identifiable Information Extraction

论文配图:GLiNER2-PII: A Multilingual Model for Personally Identifiable Information Extraction
图 1 · 摘自论文原文
  • 基于约束生成合成数据,训练0.3B参数小模型
  • 在SPY基准上达最高准确率,优于多个大模型
  • 适合需要隐私保护的多语言信息提取场景

可靠检测个人身份信息(PII)在现代数据处理系统中日益重要,但该任务仍具挑战:PII跨度多样、受地域和上下文影响,常嵌入噪声或半结构化文档中。我们提出GLiNER2-PII,一个由GLiNER2改进的0.3B参数小型模型,可在字符级粒度识别42种类型的PII实体。由于可共享标注数据稀缺且大规模收集真实PII存在隐私风险,我们构建了一个包含4,910条标注文本的多语言合成语料库,通过约束驱动生成管道,在多种语言、领域、格式和实体分布下生成多样化、真实的样本。在具有挑战性的SPY基准测试中,GLiNER2-PII在五种对比系统中取得最高的段级F1分数,包括OpenAI Privacy Filter和三个基于GLiNER的检测器。我们已将该模型公开发布于Hugging Face,以支持开放的PII检测研究与实际部署。

原文摘要 · Abstract (English)

Reliable detection of personally identifiable information (PII) is increasingly important across modern data-processing systems, yet the task remains difficult: PII spans are heterogeneous, locale-dependent, context-sensitive, and often embedded in noisy or semi-structured documents. We present GLiNER2-PII, a small 0.3B-parameter model adapted from GLiNER2 and designed to recognize a broad taxonomy of 42 PII entity types at character-span resolution. Training such systems, however, is constrained by the scarcity of shareable annotated data and the privacy risks associated with collecting real PII at scale. To address this challenge, we construct a multilingual synthetic corpus of 4,910 annotated texts using a constraint-driven generation pipeline that produces diverse, realistic examples across languages, domains, formats, and entity distributions. On the challenging SPY benchmark, GLiNER2-PII achieves the highest span-level F1 among five compared systems, including OpenAI Privacy Filter and three GLiNER-based detectors. We publicly release the model on Hugging Face to support further research and practical deployment of open PII detection systems.

隐私保护信息抽取多语言小模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。