arXiv:2506.03053cs.MAcs.AI2025-06被引 9

提出多智能体涌现行为评估框架,揭示群体互动中的安全风险

MAEBE: Multi-Agent Emergent Behavior Framework

  • 构建MAEBE框架,用双反转提问法测试智能体道德偏好
  • 单个与群体智能体的道德判断受问题表述影响显著,稳定性差
  • 群体中出现同辈压力导致共识偏移,需关注交互环境下的对齐问题

随着多智能体人工智能系统日益普及,传统针对孤立大模型的安全评估已显不足,难以捕捉新型涌现风险。本文提出多智能体涌现行为评估(MAEBE)框架,结合最大善基准(Greatest Good Benchmark)和新颖的双反转问题技术,发现:(1) 大语言模型的道德偏好,特别是对工具性伤害的判断,出人意料地脆弱,其结果随问题表述方式显著变化,无论在单个智能体还是群体中均如此;(2) 智能体群体的道德推理无法从孤立个体行为直接推断,因存在涌现的群体动态机制;(3) 特别地,群体中出现同辈压力导致行为趋同现象,即使有上级监督也难以避免,凸显了独特的安全与对齐挑战。研究强调必须在交互式多智能体环境中评估AI系统的安全性。

原文摘要 · Abstract (English)

Traditional AI safety evaluations on isolated LLMs are insufficient as multi-agent AI ensembles become prevalent, introducing novel emergent risks. This paper introduces the Multi-Agent Emergent Behavior Evaluation (MAEBE) framework to systematically assess such risks. Using MAEBE with the Greatest Good Benchmark (and a novel double-inversion question technique), we demonstrate that: (1) LLM moral preferences, particularly for Instrumental Harm, are surprisingly brittle and shift significantly with question framing, both in single agents and ensembles. (2) The moral reasoning of LLM ensembles is not directly predictable from isolated agent behavior due to emergent group dynamics. (3) Specifically, ensembles exhibit phenomena like peer pressure influencing convergence, even when guided by a supervisor, highlighting distinct safety and alignment challenges. Our findings underscore the necessity of evaluating AI systems in their interactive, multi-agent contexts.

多智能体安全评估涌现行为

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。