arXiv:2603.10012cs.CLcs.AI2026-03

军事大模型常拒答关键问题,本文提出评测方法并显著降低拒绝率。

Measuring and Eliminating Refusals in Military Large Language Models

  • 构建美军老兵参与的黄金基准数据集,首次量化军事领域拒答率。
  • 部分模型拒答率高达98.2%,软回避率达21.3%,影响作战决策。
  • 通过特定微调使回答率提升66.5点,适合军事AI系统开发者参考。

军事大语言模型需在危急时刻为作战人员提供准确信息,但当前模型因安全机制对涉及暴力、恐怖主义或军事技术的合法询问频繁拒答。本文由美国陆军及特种部队退伍军人共同构建首个此类黄金基准数据集,评估了31个公开模型和3个军事专用模型的拒答与回避率,发现拒答率最高达98.2%,软回避率在0%至21.3%之间。研究还分析了两个合成数据集与黄金数据集的相关性,并使用Heretic库对军事微调的gpt-oss-20b模型进行消融实验,结果显示回答率绝对提升66.5个百分点,但在其他军事任务上平均相对性能下降2%。最后建议采用中期及端到端后训练等深度专业化策略,以实现封闭军事模型的零拒答与最高任务准确率。

原文摘要 · Abstract (English)

Military Large Language Models (LLMs) must provide accurate information to the warfighter in time-critical and dangerous situations. However, today's LLMs are imbued with safety behaviors that cause the LLM to refuse many legitimate queries in the military domain, particularly those related to violence, terrorism, or military technology. Our gold benchmark for assessing refusal rates, which was developed by veterans of the US Army and special forces, is to our knowledge the first dataset of its kind. We present results for refusal and deflection rates on 31 public models and 3 military models. We observe hard rejection rates as high as 98.2% and soft deflection rates ranging from 0% to 21.3%. We also present results on two additional synthetic datasets and show their correlations with the gold dataset. Finally, we perform abliteration using the Heretic library on a military-tuned gpt-oss-20b model, showing an absolute increase in answer rate of 66.5 points but an average relative decrease of 2% on other military tasks. In our concluding remarks, we argue for deeper specialization, including with mid-training and end-to-end post-training, to achieve zero refusals and maximum military task accuracy for closed military models.

军事AI大模型拒答率模型优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。