打造专精网络安全的小模型,性能逼近大模型且体积更小。
Toward Cybersecurity-Expert Small Language Models
- 用专家引导+大模型生成构建高质量安全推理数据集
- 4B-20B参数模型在威胁情报与调查任务上超越多数主流模型
- 最小模型仍能达第二名,适合资源受限的实战部署
大型语言模型正改变日常应用,但其在网络安全领域的落地受限于高质量领域模型和训练数据的缺乏。为此,我们提出CyberPal 2.0,一个从4B到20B参数的网络安全专家小型语言模型系列。为训练CyberPal 2.0,我们构建了增强型链式思维安全指令数据集,基于SecKnowledge 2.0数据增强与格式化流程,融合专家介入引导推理格式与大模型驱动的多步定位,生成更高保真度、任务对齐的推理轨迹。在多种网络安全基准测试中,CyberPal 2.0持续优于基线模型,并匹配或超越多个开源与闭源前沿模型,同时仅为其极小规模。在核心网络威胁情报知识任务中,其表现几乎超越所有测试模型,仅次于Sec-Gemini v1;在核心威胁调查任务(如漏洞与缺陷工单关联弱点)中,最优20B模型超越GPT-4o、o1、o3-mini及Sec-Gemini v1,排名第一,而最小4B模型位列第二。
原文摘要 · Abstract (English)
Large language models (LLMs) are transforming everyday applications, yet deployment in cybersecurity lags due to a lack of high-quality, domain-specific models and training datasets. To address this gap, we present CyberPal 2.0, a family of cybersecurity-expert small language models (SLMs) ranging from 4B-20B parameters. To train CyberPal 2.0, we generate an enriched chain-of-thought cybersecurity instruction dataset built with our data enrichment and formatting pipeline, SecKnowledge 2.0, which integrates expert-in-the-loop steering of reasoning formats alongside LLM-driven multi-step grounding, yielding higher-fidelity, task-grounded reasoning traces for security tasks. Across diverse cybersecurity benchmarks, CyberPal 2.0 consistently outperforms its baselines and matches or surpasses various open and closed-source frontier models, while remaining a fraction of their size. On core cyber threat intelligence knowledge tasks, our models outperform almost all tested frontier models, ranking second only to Sec-Gemini v1. On core threat-investigation tasks, such as correlating vulnerabilities and bug tickets with weaknesses, our best 20B-parameter model outperforms GPT-4o, o1, o3-mini, and Sec-Gemini v1, ranking first, while our smallest 4B-parameter model ranks second.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。