构建对抗性多智能体系统测试基准,评估欺骗型智能体的隐蔽攻击与快速防御适应能力。
GAMBIT: A Three-Mode Benchmark for Adversarial Robustness in Multi-Agent LLM Collectives

- 设计三模式评测框架,涵盖零样本检测与少量样本自适应重校准。
- 生成27,804条标注数据,涵盖240种协同进化欺骗策略,实测检测率仅50.5% F1。
- 揭示零样本评估易误导:部分模型少样本适应速度差8倍,元学习者快20倍。
在多智能体系统中,单一欺骗性智能体可瓦解整个集体的收益并绕过部署防御。现有研究仅关注浅层任务且忽略自适应对手——后者会演化策略以规避检测器。为此,我们提出GAMBIT,一个包含三种评估模式的基准,配有两项独立评分指标:前两种模式评估在分布偏移增加下的零样本检测性能,第三种再校准模式衡量检测器仅用20个标注样本即适应新攻击的速度。该基准包含27,804个标注实例,覆盖240种协同进化的欺骗策略。贡献包括:(1)以国际象棋为深层推理任务,使用Gemini 3.1 Pro作为智能体,发布GAMBIT及数据集,用于在真实约束下评估对抗性检测器;(2)提出基于高效进化框架的自适应欺骗智能体,可使集体任务性能崩溃但几乎无法被检测(对Gemini检测器的F1仅为50.5%);(3)发现零样本评估对自适应对手极具误导性:两个零样本表现相近的检测器在少样本适应上相差8倍,而元学习版本收敛速度提升20倍,此差异仅在再校准模式中显现。GAMBIT是首个让攻击与防御共同演化的多智能体基准,其欺骗框架具泛化潜力,并提供快速重校准的有效方法。
原文摘要 · Abstract (English)
In multi-agent systems (MAS), a single deceptive agent can nullify all gains of an agentic AI collective and evade deployed defenses. However, existing adversarial studies on MAS target only shallow tasks and do not consider adaptive adversaries, which evolve their strategies to evade the very detectors trained to catch them. To address that gap, we introduce GAMBIT, a benchmark with three evaluation modes and two independent scores for evaluating imposter detectors: the first two modes measure zero-shot detection under increasing distribution shift, and a third recalibration mode measures how quickly a detector adapts to novel attacks from just 20 labeled examples. The benchmark comes with a dataset of 27,804 labeled instances spanning 240 co-evolved imposter strategies. Our contributions are threefold: (1) Using chess as a substrate deep reasoning problem and Gemini 3.1 Pro for agents, we release GAMBIT and its dataset to evaluate imposter detectors under realistic constraints against a stealthy adaptive imposter; (2) We introduce an adaptive imposter agent based on an efficient evolutionary framework, generalizable beyond chess, that collapses collective task performance while remaining essentially undetectable (50.5% F1-score with a Gemini-based detector); (3) We show that zero-shot evaluation can be highly misleading for adaptive adversaries: two detectors with near-identical zero-shot scores differ by 8x on few-shot adaptation, while the meta-learned variant converges 20x faster, a gap only visible in the recalibration mode. Altogether, GAMBIT provides the first multi-agent benchmark where adversarial attacks and defenses co-evolve, with an imposter framework generalizable beyond our use case, and promising techniques for fast recalibration in a rapidly evolving adversarial system. Code and data: https://anonymous.4open.science/r/gambit.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。