arXiv:2502.00757cs.CRcs.AI2025-02NeurIPS被引 9

提出自进化框架AgentBreeder,发现多智能体架构存在安全风险并可主动缓解。

AgentBreeder: Mitigating the AI Safety Risks of Multi-Agent Scaffolds via Self-Improvement

  • 通过多目标演化搜索自动优化多智能体架构设计
  • 安全基准表现提升79.4%,同时保持或提升任务能力
  • 揭示能力增强可能伴随安全漏洞,适合安全研究者参考

将大语言模型构造成多智能体系统常能提升复杂任务表现,但其安全影响尚未充分探索。本文提出AgentBreeder框架,实现对多智能体架构的多目标自进化搜索。在广泛使用的推理、数学和安全基准上评估所发现的架构,并与主流基线对比。在“蓝模式”下,安全基准平均性能提升79.4%,同时维持或提高能力得分;在“红模式”下,能力优化过程中出现对抗性脆弱的架构。本工作揭示了多智能体架构的安全风险,并提供相应缓解框架。代码已开源:https://github.com/jrosseruk/AgentBreeder。

原文摘要 · Abstract (English)

Scaffolding Large Language Models (LLMs) into multi-agent systems often improves performance on complex tasks, but the safety impact of such scaffolds has not been thoroughly explored. We introduce AgentBreeder, a framework for multi-objective self-improving evolutionary search over scaffolds. We evaluate discovered scaffolds on widely recognized reasoning, mathematics, and safety benchmarks and compare them with popular baselines. In "blue" mode, we see a 79.4% average uplift in safety benchmark performance while maintaining or improving capability scores. In "red" mode, we find adversarially weak scaffolds emerging concurrently with capability optimization. Our work demonstrates the risks of multi-agent scaffolding and provides a framework for mitigating them. Code is available at https://github.com/jrosseruk/AgentBreeder.

多智能体安全风险自进化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。