arXiv:2606.28514cs.AIcs.CL2026-06

测试多模态智能体在实时协作中应对高压与信息不对称的能力。

GPTNT: Benchmarking Real-Time Collaboration Between Multimodal Agents on Keep Talking And Nobody Explodes

论文配图:GPTNT: Benchmarking Real-Time Collaboration Between Multimodal Agents on Keep Talking And Nobody Explodes
图 1 · 摘自论文原文
  • 基于真实游戏构建实时异步协作环境,模拟真实沟通挑战。
  • 现有模型均无法在限时内成功拆弹,暴露其状态追踪与纠错短板。
  • 支持动态生成关卡,适合评估长期协作性能,适合研究人机协同。

多模态模型正越来越多地用于与人类或其他智能体协作完成任务。现有基准测试虽验证了模型的多项基础能力,但协作中的关键条件——如时间压力、信息不对称和不完整沟通——通常被孤立研究。我们提出 GPTNT,一个基于合作游戏《Keep Talking and Nobody Explodes》的基准测试,两名智能体需协作在倒计时内拆解程序化生成的炸弹谜题。一名智能体可看见并操作炸弹,但无拆弹说明;另一名拥有说明,但无法看见或操作炸弹。任一智能体都无法独立成功,必须依赖高效实时通信。与传统的回合制代理不同,GPTNT 要求异步行动与实时沟通。该基准通过隐藏说明手册或搭档,分离出模型即时推理与记忆依赖。我们发现,当前最先进系统均无法在实时中成功拆弹,而人类玩家可达成此目标。控制实验揭示其在状态跟踪、高压下效率、歧义处理与错误恢复方面的显著缺陷。GPTNT 已开源,作为当前评估未覆盖的协作性能基准。因其运行于真实游戏,支持程序化生成,并继承活跃的模组社区,可随模型进步持续演进,避免一次性被破解。

原文摘要 · Abstract (English)

Multimodal models are increasingly deployed to solve tasks collaboratively with humans or other artificial agents. Existing benchmarks show that these models possess many of the required component capabilities, but the conditions that coincide in collaboration, including time pressure, information asymmetry, and imperfect communication, are usually studied in isolation. We introduce GPTNT, a benchmark built on the cooperative video game Keep Talking and Nobody Explodes, in which two agents must coordinate to defuse procedurally generated bomb puzzles against a live countdown. One agent can see and manipulate the bomb but does not have the defusal instructions; the other has the instructions but cannot see or manipulate the bomb. Neither agent can succeed alone: success requires effective and efficient communication. Unlike turn-based proxies, GPTNT requires agents to act asynchronously and communicate in real time. GPTNT is designed to separate collaboration from reliance on memorized solutions: the instruction manual, the partner, or both can be withheld to isolate what a model derives in the moment from what it already knows. We show that GPTNT poses a substantial challenge for state-of-the-art systems: none of the closed- or open-source models we test defuses a single bomb in real time, a bar that human players clear. Through controlled experiments, we identify critical weaknesses in state tracking, efficient action under time pressure, ambiguity handling, and error recovery. We release GPTNT as a benchmark for collaborative performance that current evaluations leave unmeasured. Because it runs on the real game, GPTNT benefits from procedural generation and inherits a living modding community, allowing the benchmark to evolve as models improve rather than being solved once and retired.

多智能体协作实时通信游戏基准评估挑战

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。