arXiv:2510.03514cs.CYcs.AI2025-10被引 3

测试大模型在军事决策中的法律与道德风险,发现其行为不可预测且可能违反国际法。

Red Lines and Grey Zones in the Fog of War: Benchmarking Legal Risk, Moral Harm, and Regional Bias in Large Language Model Military Decision-Making

  • 设计四类指标,评估模型在模拟冲突中对平民目标的误伤和伤害容忍度
  • 三款主流模型均出现16.7%至66.7%的平民目标误判,伤害容忍度随局势升级上升
  • 不同模型风险差异显著,选型即意味着选择军事行动的伦理与法律风险

随着军事机构考虑将大语言模型(LLMs)用于指挥控制(C2)系统进行规划与决策支持,理解其行为倾向至关重要。本研究构建了一个评估目标行为中法律与道德风险的基准框架,通过多轮模拟冲突中让LLM作为代理进行对比测试。引入四项基于国际人道法(IHL)与军事条令的指标:平民目标率(CTR)与双重用途目标率(DTR)评估法律合规性,均值与最大模拟非战斗人员伤亡值(SNCV)量化对平民伤害的容忍度。在三个地理区域开展90次多智能体、多轮次危机模拟中,评估GPT-4o、Gemini-2.5与LLaMA-3.1三款前沿模型。结果表明,现成使用的LLM在模拟冲突中表现出令人担忧且不可预测的目标选择行为。所有模型均违反区分原则,针对平民设施的违规率介于16.7%至66.7%之间;伤害容忍度随危机推进显著上升,均值SNCV从早期回合的16.5升至晚期的27.7。模型间差异明显:LLaMA-3.1平均每次模拟执行3.47次平民打击,均值SNCV为28.4;而Gemini-2.5仅执行0.90次,均值SNCV为17.6。这些差异表明,部署模型的选择实质上是关于军事行动中可接受法律与道德风险的权衡。本研究旨在揭示在决策支持系统(AI DSS)中使用LLM可能引发的行为风险,并提供一个可复现、具有可解释性的基准测试框架,用于标准化前置测试。

原文摘要 · Abstract (English)

As military organisations consider integrating large language models (LLMs) into command and control (C2) systems for planning and decision support, understanding their behavioural tendencies is critical. This study develops a benchmarking framework for evaluating aspects of legal and moral risk in targeting behaviour by comparing LLMs acting as agents in multi-turn simulated conflict. We introduce four metrics grounded in International Humanitarian Law (IHL) and military doctrine: Civilian Target Rate (CTR) and Dual-use Target Rate (DTR) assess compliance with legal targeting principles, while Mean and Max Simulated Non-combatant Casualty Value (SNCV) quantify tolerance for civilian harm. We evaluate three frontier models, GPT-4o, Gemini-2.5, and LLaMA-3.1, through 90 multi-agent, multi-turn crisis simulations across three geographic regions. Our findings reveal that off-the-shelf LLMs exhibit concerning and unpredictable targeting behaviour in simulated conflict environments. All models violated the IHL principle of distinction by targeting civilian objects, with breach rates ranging from 16.7% to 66.7%. Harm tolerance escalated through crisis simulations with MeanSNCV increasing from 16.5 in early turns to 27.7 in late turns. Significant inter-model variation emerged: LLaMA-3.1 selected an average of 3.47 civilian strikes per simulation with MeanSNCV of 28.4, while Gemini-2.5 selected 0.90 civilian strikes with MeanSNCV of 17.6. These differences indicate that model selection for deployment constitutes a choice about acceptable legal and moral risk profiles in military operations. This work seeks to provide a proof-of-concept of potential behavioural risks that could emerge from the use of LLMs in Decision Support Systems (AI DSS) as well as a reproducible benchmarking framework with interpretable metrics for standardising pre-deployment testing.

军事决策伦理风险大模型评估国际法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。