arXiv:2606.24391cs.AIcs.CL2026-06

用1对1对抗测试大模型在信息模糊下的推理与可信度

Age of LLM: A Strategic 1v1 Benchmark for Reasoning, Diplomacy and Reliability of Large Language Models under Fog of War

论文配图:Age of LLM: A Strategic 1v1 Benchmark for Reasoning, Diplomacy and Reliability of Large Language Models under Fog of War
图 1 · 摘自论文原文
  • 两模型在隐藏地图上对战,需通过消息谈判与行动决策
  • 85%对局采用核突袭,且多为机械性发射而非战略误判
  • 模型常发无效动作,反映其状态跟踪能力不足

我们提出Age of LLM,一个基于13x7网格的回合制1对1对抗基准,两大型语言模型需摧毁对方基地。设计三大压力机制:战场迷雾、完整外交(含消息、停火、最后通牒;铀资源保密)及可靠性约束(每步必须符合严格JSON格式,违规操作将被静默丢弃)。引擎私有,每局使用随机地图与对手,避免公开基准的数据污染问题。模型仅获近似规则提示,无建造顺序建议(数据收集期间仅存在两个战术引导词,见第2.7节)。在54场对局中评测15个推理模型,共5,258次行动。发现:(1)核突袭主导局势(规则一致v0.11+子语料中占78%,全语料85%),其单一发射特征在秘密同时发射规则下主要为机械行为,非认知威慑失效;(2)军事征服虽罕见但更快(平均12.3轮对比18.9轮);(3)外交频繁但极少达成协议;(4)约58%的非法操作源于迷雾或状态错误,使其成为信念追踪能力的衡量指标;(5)首次观察到可靠性与胜率弱关联,尚属探索性发现。语料规模小、不平衡且未左右互换,排名仅为初步描述,非核心贡献。除排名外,逐回合动作与对话记录为研究模型在对抗不确定性下的推理过程提供了窗口——包括信念追踪、自发欺骗与模型级认知‘人格’,视为未来方向。已发布回放格式、等距视角可视化工具及全部回放;引擎源码可申请获取。

原文摘要 · Abstract (English)

We introduce Age of LLM, a turn-based 1v1 benchmark in which two LLMs face off on a 13x7 grid to destroy the enemy base. Three stressors are deliberate: fog of war, full diplomacy (messages, ceasefires, ultimatums; uranium kept secret), and a reliability dimension where every turn must follow a strict JSON schema and an illegal action is silently discarded. The engine is private and each match uses a fresh random map seed and opponent, mitigating the data contamination that affects public benchmarks. Models receive a (near) rule-only prompt with no build-order advice (two tactical seed phrases were present during data collection; see Section 2.7). We benchmark 15 reasoning models across 54 matches and 5,258 actions. Findings: (1) the nuclear rush dominates (78% on the rules-coherent v0.11+ sub-corpus; 85% corpus-wide) with a sole-launcher signature that is largely mechanical under secret-simultaneous launch rules, not a cognitive deterrence failure; (2) military conquest is rare but faster (12.3 vs 18.9 turns); (3) diplomacy is prolific yet almost never consummated; (4) ~58% of illegal actions are fog/state errors, making the illegal-action rate a measure of belief-tracking; (5) -- the least established, and the only one we label exploratory -- a weak link associates reliability with winning. The corpus is small, unbalanced and not side-swapped, so the ranking is a preliminary descriptive view, not a contribution. Beyond ranking, the turn-by-turn traces of actions and messages make the corpus a lens on how LLMs reason under adversarial uncertainty -- their belief-tracking, spontaneous deception, and per-model cognitive "personas" -- which we frame as a future research direction. We release the replay format, an isometric viewer and all replays; engine source on request.

对抗推理大模型评估信念追踪游戏基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。