arXiv:2510.26550cs.AI2025-10被引 1

200亿参数模型在边缘设备上实现军事任务媲美GPT-5

EdgeRunner 20B: Military Task Parity with GPT-5 while Running on the Edge

  • 基于160万条军用资料微调,专攻军事场景
  • 军事测试集上95%以上显著优于或持平GPT-5
  • 本地部署安全高效,适合离线敏感环境

我们提出EdgeRunner 20B,是针对军事任务优化的gpt-oss-20b微调版本。该模型在160万条精选军用文档与网站数据上训练。我们还构建了四个新测试集:(a) 作战兵种,(b) 作战医疗,(c) 网络行动,(d) mil-bench-5k(通用军事知识)。在这些军事测试集上,EdgeRunner 20B在多数任务中达到或超越GPT-5表现,统计显著性达95%以上,仅在作战医疗高推理和mil-bench-5k低推理设置下未达标。与gpt-oss-20b相比,在通用基准如ARC-C、GPQA Diamond、GSM8k、IFEval、MMLU Pro、TruthfulQA上无显著退化,仅GSM8k在低推理设置下有轻微下降。我们还分析了超参数、成本与吞吐量。结果表明,小型本地化模型是军事等数据敏感场景的理想方案,可部署于隔离的边缘设备。

原文摘要 · Abstract (English)

We present EdgeRunner 20B, a fine-tuned version of gpt-oss-20b optimized for military tasks. EdgeRunner 20B was trained on 1.6M high-quality records curated from military documentation and websites. We also present four new tests sets: (a) combat arms, (b) combat medic, (c) cyber operations, and (d) mil-bench-5k (general military knowledge). On these military test sets, EdgeRunner 20B matches or exceeds GPT-5 task performance with 95%+ statistical significance, except for the high reasoning setting on the combat medic test set and the low reasoning setting on the mil-bench-5k test set. Versus gpt-oss-20b, there is no statistically-significant regression on general-purpose benchmarks like ARC-C, GPQA Diamond, GSM8k, IFEval, MMLU Pro, or TruthfulQA, except for GSM8k in the low reasoning setting. We also present analyses on hyperparameter settings, cost, and throughput. These findings show that small, locally-hosted models are ideal solutions for data-sensitive operations such as in the military domain, allowing for deployment in air-gapped edge devices.

边缘计算军事AI大模型微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。