arXiv:2508.01056cs.ETcs.AI2025-08被引 1

用简单方法控制大模型在军事场景中的升级倾向。

Managing Escalation in Off-the-Shelf Large Language Models

  • 通过非技术干预降低模型在地缘博弈中的激进建议。
  • 实验中显著减少战略推演游戏中的升级行为。
  • 适合关注AI安全与政策应用的决策者参考。

美国国家安全相关方已开始使用大型语言模型,包括公众熟知的企业版“现成”模型(如ChatGPT)。这一趋势将加速。然而,近期研究表明,这类模型在面对地缘政治或战略情景时,常会提出激化行动。本文展示两种简单、非技术性干预措施,可有效抑制此类倾向。将这些干预纳入近期一项研究的对抗推演设计后,游戏全程的升级行为大幅减少。因此,限制大模型在国家安全领域应用的呼吁尚不成熟。美国政府已并将继续利用大模型进行情景规划与行动方案建议。本研究承认大模型的必然应用,提供可操作措施以使其符合国家安全目标,尤其是升级管理。

原文摘要 · Abstract (English)

U.S. national security customers have begun to utilize large language models, including enterprise versions of ``off-the-shelf'' models (e.g., ChatGPT) familiar to the public. This uptake will likely accelerate. However, recent studies suggest that off-the-shelf large language models frequently suggest escalatory actions when prompted with geopolitical or strategic scenarios. We demonstrate two simple, non-technical interventions to control these tendencies. Introducing these interventions into the experimental wargame design of a recent study, we substantially reduce escalation throughout the game. Calls to restrict the use of large language models in national security applications are thus premature. The U.S. government is already, and will continue, employing large language models for scenario planning and suggesting courses of action. Rather than warning against such applications, this study acknowledges the imminent adoption of large language models, and provides actionable measures to align them with national security goals, including escalation management.

大模型安全军事推演策略控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。