提出路由专家框架,让安全防护模型更精准识别不同威胁。
Triaging Threats to Specialized Guardrails

- 用路由机制将对话分配给专门的防护专家模型
- 在32,460样本上测试,比单一模型提升细粒度检测能力
- 适合需要灵活扩展新威胁的生产级应用
构建鲁棒的安全防护机制对大语言模型在多样化实际场景中的部署至关重要。然而,这一目标面临挑战:安全风险涵盖异构威胁领域,而现有数据集仅覆盖零散的风险子集,且分类体系不一致。因此,当前防护模型能否泛化到狭窄评估之外仍不清楚。为更深入理解防护模型的鲁棒性,我们首次提出GuardZoo——一个统一的人工标注基准,包含32,460个样本,覆盖15类不同危险类别。在GuardZoo上的评估表明,单体式防护模型存在任务干扰问题:不同威胁领域需要不同的决策边界,难以压缩进单一模型。为此,我们提出RouteGuard,一种路由-专家框架,可将每轮对话分发至特定领域的专家防护模型进行威胁检测。实验显示,RouteGuard在强基线基础上显著提升细粒度威胁检测性能,跨域泛化能力更强,并支持对新兴威胁的灵活模块化扩展。
原文摘要 · Abstract (English)
Building robust safety guardrails is essential for deploying Large Language Models across diverse real-world applications. However, this goal remains challenging because safety risks span heterogeneous threat domains, while existing datasets cover only fragmented risk subsets and rely on inconsistent taxonomies. Consequently, it remains unclear whether current guardrails can generalize beyond narrow evaluation settings. To better understand the robustness of guardrail models, we first introduce GuardZoo, a unified human-annotated benchmark with 32,460 samples covering 15 distinct unsafe categories. Evaluation on GuardZoo reveals that monolithic guardrails suffer from task interference: different threat domains require distinct decision boundaries that are difficult to compress into a single model. We therefore propose RouteGuard, a router-expert framework that triages each conversation to specialized expert guardrails for threat-specific detection. Experiments show that RouteGuard improves fine-grained threat detection over strong guardrail baselines, generalizes better under out-of-domain evaluation, and supports flexible modular expansion to emerging threats.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。