为美军场景定制大模型,提升军事术语理解能力。
Fine-Tuning and Evaluating Open-Source Large Language Models for the Army Domain
- 用军方语料微调开源大模型,增强领域适应性。
- 三轮迭代后,模型在军务任务上表现持续提升。
- 开发评估框架MilBench,量化模型军事知识掌握度。
近年来,大型语言模型(LLMs)在军事领域的应用引发关注。然而,现有模型在军用场景中表现不佳,主要因缺乏专业术语和行话。为克服这一问题,许多机构选择微调而非从头训练新模型。本文介绍由陆军未来司令部(AFC)下属研究与分析中心(TRAC)开发的三代TRACLM系列模型,通过持续优化训练流程,每轮迭代均显著提升模型在军务任务中的表现。同时,为客观评估模型的军事知识水平,我们构建了可扩展的评估框架MilBench,基于条令和考核任务对模型进行测试。本文公开初步结果、模型、方法及建议,为国防部内大模型技术发展提供重要参考,支持高层人工智能整合决策。
原文摘要 · Abstract (English)
In recent years, the widespread adoption of Large Language Models (LLMs) has sparked interest in their potential for application within the military domain. However, the current generation of LLMs demonstrate sub-optimal performance on Army use cases, due to the prevalence of domain-specific vocabulary and jargon. In order to fully leverage LLMs in-domain, many organizations have turned to fine-tuning to circumvent the prohibitive costs involved in training new LLMs from scratch. In light of this trend, we explore the viability of adapting open-source LLMs for usage in the Army domain in order to address their existing lack of domain-specificity. Our investigations have resulted in the creation of three distinct generations of TRACLM, a family of LLMs fine-tuned by The Research and Analysis Center (TRAC), Army Futures Command (AFC). Through continuous refinement of our training pipeline, each successive iteration of TRACLM displayed improved capabilities when applied to Army tasks and use cases. Furthermore, throughout our fine-tuning experiments, we recognized the need for an evaluation framework that objectively quantifies the Army domain-specific knowledge of LLMs. To address this, we developed MilBench, an extensible software framework that efficiently evaluates the Army knowledge of a given LLM using tasks derived from doctrine and assessments. We share preliminary results, models, methods, and recommendations on the creation of TRACLM and MilBench. Our work significantly informs the development of LLM technology across the DoD and augments senior leader decisions with respect to artificial intelligence integration.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。