arXiv:2603.16642cs.AIcs.CL2026-03

用未训练过的2026中东冲突数据,测试AI在战争迷雾中推理真实战略的能力。

When AI Navigates the Fog of War

  • 构建11个时间节点与42个可验证问题,确保AI仅基于当时公开信息推理。
  • 顶尖大模型能洞察深层结构动因,但在多势力政治环境中表现不稳定。
  • 模型判断随时间演变,从速战速决转向对长期僵局的系统性分析。

当前大语言模型能否在战争轨迹尚未明朗时进行合理推断?由于事后预测易受训练数据泄露干扰,这一问题难以评估。我们通过一个时间上严谨的案例研究,分析2026年中东冲突早期阶段——该事件发生在当前前沿模型的训练截止之后。我们设立了11个关键时间节点、42个针对各节点的可验证问题以及5个通用探索性问题,要求模型仅依据每个时刻可公开获取的信息进行推理。这一设计显著缓解了训练数据泄露问题,为研究模型在“战争迷雾”中的推理能力提供了新场景,并据我们所知,首次实现了对大模型在持续进行的地缘政治冲突中推理行为的时间化分析。结果显示:第一,当前最先进模型常表现出显著的战略现实感,能超越表面言论,洞察深层结构性动因;第二,该能力在经济与后勤结构清晰的领域更可靠,在政治模糊的多方博弈环境中则表现不均;第三,模型叙事随时间演进,从初期预期快速控制,转向后期对区域固化与消耗式缓和的系统性理解。由于冲突仍在持续,本研究可作为模型在危机发展过程中的档案快照,为未来研究提供无回溯偏见的分析基准。

原文摘要 · Abstract (English)

Can AI reason about a war before its trajectory becomes historically obvious? Analyzing this capability is difficult because retrospective geopolitical prediction is heavily confounded by training-data leakage. We address this challenge through a temporally grounded case study of the early stages of the 2026 Middle East conflict, which unfolded after the training cutoff of current frontier models. We construct 11 critical temporal nodes, 42 node-specific verifiable questions, and 5 general exploratory questions, requiring models to reason only from information that would have been publicly available at each moment. This design substantially mitigates training-data leakage concerns, creating a setting well-suited for studying how models analyze an unfolding crisis under the fog of war, and provides, to our knowledge, the first temporally grounded analysis of LLM reasoning in an ongoing geopolitical conflict. Our analysis reveals three main findings. First, current state-of-the-art large language models often display a striking degree of strategic realism, reasoning beyond surface rhetoric toward deeper structural incentives. Second, this capability is uneven across domains: models are more reliable in economically and logistically structured settings than in politically ambiguous multi-actor environments. Finally, model narratives evolve over time, shifting from early expectations of rapid containment toward more systemic accounts of regional entrenchment and attritional de-escalation. Since the conflict remains ongoing at the time of writing, this work can serve as an archival snapshot of model reasoning during an unfolding geopolitical crisis, enabling future studies without the hindsight bias of retrospective analysis.

战略推理地缘政治大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。