用极少数据让大模型精通网络安全,训练量减少42倍仍超顶尖模型。
Less Data, More Security: Advancing Cybersecurity LLMs Specialization via Resource-Efficient Domain-Adaptive Continuous Pre-training with Minimal Tokens
- 通过分布式训练在少量安全语料上持续微调大模型。
- 700亿参数模型仅用1.188亿词元,准确率超现有顶尖安全模型。
- 适合需要高效、低能耗安全AI助手的团队使用。
随着人工智能工作负载规模增长,亟需可扩展且可持续的高性能计算基础设施与训练方法。尽管大语言模型在自然语言处理上表现卓越,但通用模型往往缺乏有效网络安全分析所需的领域知识。本文研究了领域自适应连续预训练(DAP)作为一种可扩展、资源高效的增强方式,通过多节点GPU集群上的全分片数据并行(FSDP)流水线实现。我们系统地对三种解码器架构——Llama-3.1-8B、DeepSeek-R1-Distill-Qwen-14B 和 Llama-3.3-70B-Instruct——进行优化,使用来自标准、学术文献和技术文档的1.26亿词元网络安全语料库。在三个网络安全基准测试(CTI-MCQ、CyberMetric、SecEval)中,适配后性能显著提升。特别地,我们的 Llama-3.3-70B-Ins-DAP 模型分别取得0.718、0.933和0.864的准确率,超越参数高效基线与专用模型(如 Llama-Primus-Base,训练数据27.7亿词元;Foundation-Sec-8B,50亿词元),且仅使用1.188亿词元,相当于训练数据量减少23至42倍。基于可扩展HPC基础设施的定向连续预训练,实现了低计算与能耗的网络安全领域适配,支持威胁分析、漏洞评估与安全文档生成等专业应用,推动可持续负责任的人工智能发展。
原文摘要 · Abstract (English)
The increasing scale of AI workloads demands High-Performance Computing (HPC) infrastructure and training methodologies that are both scalable and sustainable. While Large Language Models (LLMs) demonstrate exceptional natural language capabilities, general-purpose models often lack the specialized domain knowledge necessary for effective cybersecurity analysis. We investigate Domain-Adaptive Continuous Pretraining (DAP) as a scalable, resource-efficient methodology for enhancing cybersecurity understanding in pretrained LLMs, implemented through a distributed Fully Sharded Data Parallel (FSDP) pipeline across multi-node GPU clusters. We systematically adapted three decoder-based architectures -- Llama-3.1-8B, DeepSeek-R1-Distill-Qwen-14B, and Llama-3.3-70B-Instruct -- using a curated 126-million-word cybersecurity corpus from standards, academic literature, and technical documentation. Evaluation across three cybersecurity benchmarks -- CTI-MCQ, CyberMetric, and SecEval -- demonstrates consistent improvements post-adaptation. Notably, our Llama-3.3-70B-Ins-DAP model achieves state-of-the-art performance with accuracies of 0.718, 0.933, and 0.864, respectively, surpassing parameter-efficient baselines and specialized models including Llama-Primus-Base (trained on 2.77 billion tokens) and Foundation-Sec-8B (trained on 5 billion tokens), despite utilizing only 118.8 million tokens -- representing a 23-to-42-fold reduction in training data. Targeted continuous pretraining via scalable HPC infrastructure enables effective cybersecurity domain adaptation with a substantially reduced computational and energy footprint, supporting specialized AI assistants in threat analysis, vulnerability assessment, and security documentation, while advancing sustainable and responsible AI development.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。