测试大模型在军事危机中的自由回答不一致性,发现其决策极不稳定。
Measuring Free-Form Decision-Making Inconsistency of Language Models in Military Crisis Simulations
- 用BERTScore量化模型自由回答的语义不一致程度
- 5个模型在不同条件下均显示显著决策不一致
- 提示词微调比温度变化更易引发决策分歧,适合高风险场景评估
随着各国探索使用语言模型辅助军事危机决策,本文在类美军推演的危机模拟中,考察了模型自由回答的不一致性。以往研究受限于预设动作集,难以量化自然语言决策差异。本文采用基于BERTScore的度量方法,有效捕捉语义不变下的语言变体影响。实验表明,五个测试模型在调整推演设定、匿名化国家或改变采样温度T时,仍表现出显著的语义不一致。定性分析显示,模型推荐的行动方案相似度极低。进一步研究发现,在T=0时,语义等价的提示词变化导致的不一致性,普遍超过温度采样带来的差异。鉴于军事决策的高风险性,建议谨慎使用语言模型支持重大决策。
原文摘要 · Abstract (English)
There is an increasing interest in using language models (LMs) for automated decision-making, with multiple countries actively testing LMs to aid in military crisis decision-making. To scrutinize relying on LM decision-making in high-stakes settings, we examine the inconsistency of responses in a crisis simulation ("wargame"), similar to reported tests conducted by the US military. Prior work illustrated escalatory tendencies and varying levels of aggression among LMs but were constrained to simulations with pre-defined actions. This was due to the challenges associated with quantitatively measuring semantic differences and evaluating natural language decision-making without relying on pre-defined actions. In this work, we query LMs for free form responses and use a metric based on BERTScore to measure response inconsistency quantitatively. Leveraging the benefits of BERTScore, we show that the inconsistency metric is robust to linguistic variations that preserve semantic meaning in a question-answering setting across text lengths. We show that all five tested LMs exhibit levels of inconsistency that indicate semantic differences, even when adjusting the wargame setting, anonymizing involved conflict countries, or adjusting the sampling temperature parameter $T$. Further qualitative evaluation shows that models recommend courses of action that share few to no similarities. We also study the impact of different prompt sensitivity variations on inconsistency at temperature $T = 0$. We find that inconsistency due to semantically equivalent prompt variations can exceed response inconsistency from temperature sampling for most studied models across different levels of ablations. Given the high-stakes nature of military deployment, we recommend further consideration be taken before using LMs to inform military decisions or other cases of high-stakes decision-making.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。