arXiv:2509.17192cs.AI2025-09被引 1

9%的策略模拟研究让大模型同时控制行动与裁决,揭示了其在开放对战中的角色局限。

Shall We Play a Game? Language Models for Open-ended Wargames

  • 分析223篇论文,按模型控制角色分类:行动、裁决或两者兼有。
  • 仅20篇(约9%)研究让大模型同时负责玩家行为和结果判定。
  • 强调开放对抗中需明确模型控制范围,否则仿真可信度存疑。

基于大语言模型的社会模拟可使生成对话看似单一行为信号,但模型背后可能承担多重任务:决定角色言行、判断行动后果,或两者兼具。这一差异在开放性兵棋推演中尤为关键,因模型需处理非常规动作与模糊结果。我们对截至2026年5月1日检索到的223篇去重的AI参与兵棋推演与战略模拟论文进行了初步综述,依据每项模拟中模型的控制属性进行分类:是否对玩家行为、裁决过程或两者均具有开放控制权。结果显示,仅有20篇(约9%)研究赋予大模型双重角色。在将语言模型输出视为社会模拟前,研究者必须明确模型对行为与后果的自主控制程度。对于开放性模拟而言,仿真真实感不仅取决于代理行为是否合理,更取决于语言模型能否稳定充当裁判或世界模型。

原文摘要 · Abstract (English)

LLM-based social simulations can make a generated transcript look like a single behavioral signal, but the model behind that transcript may be doing several different jobs: choosing what an actor says or does, deciding what happens after an action, or both. The difference matters especially in open-ended wargames, where models are prized for handling unusual actions and ambiguous consequences. We report a scoping review of 223 de-duplicated AI-in-wargames and strategic-simulation papers retrieved through May 1, 2026, describing each simulation by its model-control profile: whether the language model has open-ended control over player actions, adjudication, or both. Only 20 of 223 studies (~9%) give language models both roles. Before treating LM outputs as social simulations, researchers need to know how much creative control the model has over actions and consequences. For open-ended simulations, fidelity depends not only on whether agents behave plausibly, but also on whether language models can reliably act as adjudicators or world models.

兵棋推演大模型控制社会模拟

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。