为军事决策场景设计新基准,发现大模型在真实战场下严重失效。
WARBENCH: A Comprehensive Benchmark for Evaluating LLMs in Military Decision-Making
- 构建四维压力测试框架,涵盖法律约束、边缘计算等短板
- 小模型法律违规率近70%,4比特量化导致性能崩溃
- 显式推理机制可有效防止误判,适合高风险场景应用
大型语言模型正被考虑用于安全关键的军事应用,但现有评估基准存在结构性盲区,系统性高估模型在真实战术场景中的能力。现有框架通常忽略基于《国际人道法》(IHL)的严格法律约束,遗漏边缘计算限制,缺乏对“战争迷雾”的鲁棒性测试,且未能充分评估显式推理能力。为此,我们提出WARBENCH,一个综合性评估框架,建立基础战术基准,并包含四个独立的压力测试维度。通过对九个领先模型在136个高保真历史场景下的大规模实证评估,发现严重结构性缺陷:首先,复杂地形与高兵力不对称下,基础战术推理系统性崩溃;其次,虽顶尖闭源模型保持功能合规,但边缘优化的小模型暴露极端操作风险,法律违规率接近70%;此外,模型在4比特量化下出现灾难性性能下降与系统性信息丢失。相反,显式推理机制能有效作为结构化防护屏障,避免无意违规。最终结果表明,当前模型仍远未准备好在高风险战术环境中实现自主部署。
原文摘要 · Abstract (English)
Large Language Models are increasingly being considered for deployment in safety-critical military applications. However, current benchmarks suffer from structural blindspots that systematically overestimate model capabilities in real-world tactical scenarios. Existing frameworks typically ignore strict legal constraints based on International Humanitarian Law (IHL), omit edge computing limitations, lack robustness testing for fog of war, and inadequately evaluate explicit reasoning. To address these vulnerabilities, we present WARBENCH, a comprehensive evaluation framework establishing a foundational tactical baseline alongside four distinct stress testing dimensions. Through a large scale empirical evaluation of nine leading models on 136 high-fidelity historical scenarios, we reveal severe structural flaws. First, baseline tactical reasoning systematically collapses under complex terrain and high force asymmetry. Second, while state of the art closed source models maintain functional compliance, edge-optimized small models expose extreme operational risks with legal violation rates approaching 70 percent. Furthermore, models experience catastrophic performance degradation under 4-bit quantization and systematic information loss. Conversely, explicit reasoning mechanisms serve as highly effective structural safeguards against inadvertent violations. Ultimately, these findings demonstrate that current models remain fundamentally unready for autonomous deployment in high stakes tactical environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。