用户自定义政策的动态安全防护模型,检测更准更快。
DynaGuard: A Dynamic Guardian Model With User-Defined Policies
- 根据用户自定义政策动态判断文本风险,灵活应对新威胁。
- 在传统安全类别上准确率超静态模型,自由格式违规检测媲美顶尖推理模型。
- 支持链式思考解释输出,适合需要可解释性的高风险场景。
守护模型在确保面向用户的AI应用安全与伦理行为方面至关重要,能施加约束并检测有害内容。现有标准守护模型仅支持预设的静态危害类别,我们提出DynaGuard——一套可基于用户自定义策略动态评估文本的守护模型,并配套DynaBench数据集用于训练与评估。该模型不仅能快速检测政策违规,还支持链式思考推理以阐明和证明输出。关键的是,DynaGuard在传统安全类别上的检测准确率超越静态模型,且在自由格式政策违规检测上达到前沿推理模型水平,耗时仅为后者的一小部分。这使其成为语言模型安全护栏的关键工具。
原文摘要 · Abstract (English)
Guardian models play a crucial role in ensuring the safety and ethical behavior of user-facing AI applications by enforcing guardrails and detecting harmful content. While standard guardian models are limited to predefined, static harm categories, we introduce DynaGuard, a suite of dynamic guardian models offering novel flexibility by evaluating text based on user-defined policies, and DynaBench, a dataset for training and evaluating dynamic guardian models. Our models provide both rapid detection of policy violations and a chain-of-thought reasoning option that articulate and justify model outputs. Critically, DynaGuard not only surpasses static models in detection accuracy on traditional safety categories, but is competitive with frontier reasoning models on free-form policy violations, all in a fraction of the time. This makes DynaGuard an critical tool for language model guardrails.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。