用多智能体辩论提升大模型安全评估效率,成本更低。
Efficient LLM Safety Evaluation through Multi-Agent Debate
- 设计多智能体框架,让批评者、辩护者和裁判协同辩论
- 在1.1万条数据上表现优于小模型基线,接近GPT-4o效果
- 只需少量辩论轮次即可实现主要收益,适合大规模应用
大型语言模型(LLM)的安全评估日益依赖以LLM为裁判的流水线,但高性能裁判成本高昂。本文研究结构化多智能体辩论能否在保持模型规模与成本适中的前提下提升裁判可靠性。为此,提出HAJailBench,一个包含11,100条人工标注交互的越狱攻击基准,涵盖多样攻击方法与目标模型。结合多智能体裁判框架,批评者、辩护者与裁判在统一安全标准下展开辩论。在HAJailBench上,该框架优于同规模小模型提示基线及已有多智能体裁判,且在所测价格下比GPT-4o更经济。消融实验表明,少量辩论轮次即可捕获大部分性能提升。结果支持结构化、价值对齐的辩论作为可扩展大模型安全评估的实际方案。
原文摘要 · Abstract (English)
Safety evaluation of large language models (LLMs) increasingly relies on LLM-as-a-judge pipelines, but strong judges can still be expensive to use at scale. We study whether structured multi-agent debate can improve judge reliability while keeping backbone size and cost modest. To do so, we introduce HAJailBench, a human-annotated jailbreak benchmark with 11,100 labeled interactions spanning diverse attack methods and target models, and we pair it with a Multi-Agent Judge framework in which critic, defender, and judge agents debate under a shared safety rubric. On HAJailBench, the framework improves over matched small-model prompt baselines and prior multi-agent judges, while remaining more economical than GPT-4o under the evaluated pricing snapshot. Ablation results further show that a small number of debate rounds is sufficient to capture most of the gain. Together, these results support structured, value-aligned debate as a practical design for scalable LLM safety evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。