用多个AI辩论自动发现并修复大模型的不安全行为。
RedDebate: Safer Responses Through Multi-Agent Red Teaming Debates
- 多智能体辩论机制自动识别模型漏洞。
- 加入长期记忆后,不安全输出减少37%以上。
- 适合关注AI安全与自动化评测的研究者。
我们提出RedDebate,一种全新的多智能体辩论框架,使大语言模型(LLMs)能够自主识别并缓解其不安全行为。现有AI安全方法多依赖昂贵的人工评估或孤立的单模型检测,存在可扩展性差和遗漏风险。RedDebate通过在多样化辩论场景中让多个LLMs协作论证,实现对彼此推理的批判性评估,从而以完全自动化方式开展红队测试,系统性挖掘不安全缺陷。为此,我们设计了独特的长期记忆模块,保留辩论中的安全相关见解,并在后续推理中复用,实现行为持续优化。在多种模型和安全基准上的实证表明,RedDebate显著降低不安全输出;仅靠辩论即可改进行为,而引入记忆后错误率进一步下降。据我们所知,RedDebate是首个完全自动化融合多智能体辩论与红队测试的框架,可无须人工干预地逐步提升LLM安全性。
原文摘要 · Abstract (English)
We introduce RedDebate, a novel multi-agent debate framework that provides the foundation for Large Language Models (LLMs) to identify and mitigate their unsafe behaviours. AI safety approaches often rely on costly human evaluation or isolated single-model assessment, both constrained by scalability and prone to oversight failures. RedDebate employs collaborative argumentation among multiple LLMs across diverse debate scenarios, enabling them to critically evaluate one another's reasoning and systematically uncover unsafe failure modes through fully automated red-teaming. To support this, we propose designing distinct long-term memory modules that preserve safety-relevant insights from debate interactions and leverage them during subsequent inference, facilitating continuous refinement of model behaviour. Empirical evaluation on safety benchmarks across a diverse set of models demonstrates that RedDebate substantially reduces unsafe outputs. While debate alone allows LLMs to refine their behaviour, the addition of memory yields further error reductions. To the best of our knowledge, RedDebate is the first fully automated framework to unify multi-agent debate and red-teaming to progressively enhance LLM safety without human intervention.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。