arXiv:2501.12456cs.CRcs.AI2025-01中稿 · AAAI被引 8

OneShield框架可有效识别多语言敏感信息,助力企业与开源项目合规

Deploying Privacy Guardrails for LLMs: A Comparative Analysis of Real-World Applications

  • 基于上下文感知的实体识别技术,动态检测输入输出中的隐私数据
  • 多语言场景下敏感信息检测F1达0.95,比现有工具高12%
  • 开源部署中三月节省超300小时人工,精准标记8.25%潜在风险代码

大型语言模型(LLMs)的广泛应用带来了显著的隐私挑战。本文详细研究OneShield Privacy Guard框架,该框架旨在缓解企业与开源环境中用户输入和模型输出的隐私风险。通过两个真实部署案例验证:(1) 集成于Data and Model Factory的多语言隐私保护系统,在26种语言上实现0.95 F1分数,优于StarPII和Presidio等前沿工具最多12%;(2) 开源项目PR Insights,平均F1为0.86,三个月内减少超过300小时人工工作量,准确识别出1,256个拉取请求中8.25%的隐私风险,且具备更强上下文敏感性。结果表明,OneShield在多种环境下均具高度适应性与有效性,为上下文感知实体识别、自动化合规及伦理AI落地提供实践路径。

原文摘要 · Abstract (English)

The adoption of Large Language Models (LLMs) has revolutionized AI applications but poses significant challenges in safeguarding user privacy. Ensuring compliance with privacy regulations such as GDPR and CCPA while addressing nuanced privacy risks requires robust and scalable frameworks. This paper presents a detailed study of OneShield Privacy Guard, a framework designed to mitigate privacy risks in user inputs and LLM outputs across enterprise and open-source settings. We analyze two real-world deployments:(1) a multilingual privacy-preserving system integrated with Data and Model Factory, focusing on enterprise-scale data governance; and (2) PR Insights, an open-source repository emphasizing automated triaging and community-driven refinements. In Deployment 1, OneShield achieved a 0.95 F1 score in detecting sensitive entities like dates, names, and phone numbers across 26 languages, outperforming state-of-the-art tool such as StarPII and Presidio by up to 12\%. Deployment 2, with an average F1 score of 0.86, reduced manual effort by over 300 hours in three months, accurately flagging 8.25\% of 1,256 pull requests for privacy risks with enhanced context sensitivity. These results demonstrate OneShield's adaptability and efficacy in diverse environments, offering actionable insights for context-aware entity recognition, automated compliance, and ethical AI adoption. This work advances privacy-preserving frameworks, supporting user trust and compliance across operational contexts.

隐私保护LLM安全实体识别合规框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。