OneShield为大模型提供可定制的安全防护,针对不同用户需求灵活设置风险控制策略。
OneShield -- the Next Generation of LLM Guardrails
- 设计了独立于模型的可定制安全框架,支持自定义风险因素与合规政策
- 部署后已支持多客户场景,具备良好可扩展性与实际应用效果
- 适合关注大模型安全合规的企业或开发者使用
大型语言模型的兴起带来了广泛应用的巨大潜力,但随之而来的安全性、隐私和伦理问题也日益突出。各方正为自身模型及独立解决方案制定保护措施。由于大模型持续演进,通用防护难以实现,统一方案不可行。本文提出 OneShield——一种独立、模型无关且可定制的安全防护方案,旨在为特定客户提供风险识别、安全策略表达与声明、风险缓解等功能。文中描述了框架实现,讨论了可扩展性,并提供了上线以来的使用统计数据。
原文摘要 · Abstract (English)
The rise of Large Language Models has created a general excitement about the great potential for a myriad of applications. While LLMs offer many possibilities, questions about safety, privacy, and ethics have emerged, and all the key actors are working to address these issues with protective measures for their own models and standalone solutions. The constantly evolving nature of LLMs makes it extremely challenging to universally shield users against their potential risks, and one-size-fits-all solutions are unfeasible. In this work, we propose OneShield, our stand-alone, model-agnostic and customizable solution to safeguard LLMs. OneShield aims to provide facilities for defining risk factors, expressing and declaring contextual safety and compliance policies, and mitigating LLM risks, with a focus on each specific customer. We describe the implementation of the framework, discuss scalability considerations, and provide usage statistics of OneShield since its initial deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。