让大模型安全评估像法庭辩论一样,无需重训就能适应新规则。
CourtGuard: A Model-Agnostic Framework for Zero-Shot Policy Adaptation in LLM Safety
- 用外部政策文档驱动多智能体对抗辩论,实现安全判断
- 在7个基准上超越现有方法,零样本适配维基破坏检测达90%准确率
- 适合需要快速响应新规或自动生成安全数据的研究者
当前大模型安全机制依赖静态微调分类器,难以在不重新训练的情况下应对新治理规则。为此,我们提出CourtGuard,一种检索增强的多智能体框架,将安全评估重构为基于证据的辩论。通过调用外部政策文档进行对抗性辩论,CourtGuard在7个安全基准上达到领先性能,且无需微调即可超越专用政策遵循基线。此外,该框架展现出两项关键能力:(1)零样本可迁移性,在更换参考政策后成功应用于跨领域的维基破坏检测任务,准确率达90%;(2)自动化数据构建与审计能力,利用CourtGuard构建并审核了九个新型复杂对抗攻击数据集。结果表明,将安全逻辑与模型权重解耦,为满足当前及未来AI治理监管要求提供了鲁棒、可解释且可扩展的路径。
原文摘要 · Abstract (English)
Current safety mechanisms for Large Language Models (LLMs) rely heavily on static, fine-tuned classifiers that suffer from adaptation rigidity, the inability to enforce new governance rules without expensive retraining. To address this, we introduce CourtGuard, a retrieval-augmented multi-agent framework that reimagines safety evaluation as Evidentiary Debate. By orchestrating an adversarial debate grounded in external policy documents, CourtGuard achieves state-of-the-art performance across 7 safety benchmarks, outperforming dedicated policy-following baselines without fine-tuning. Beyond standard metrics, we highlight two critical capabilities: (1) Zero-Shot Adaptability, where our framework successfully generalized to an out-of-domain Wikipedia Vandalism task (achieving 90\% accuracy) by swapping the reference policy; and (2) Automated Data Curation and Auditing, where we leveraged CourtGuard to curate and audit nine novel datasets of sophisticated adversarial attacks. Our results demonstrate that decoupling safety logic from model weights offers a robust, interpretable, and adaptable path for meeting current and future regulatory requirements in AI governance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。