arXiv:2502.11191cs.CRcs.AI2025-02EMNLP被引 26

发布首个开源网络安全大模型训练数据集,提升模型安全能力。

Primus: A Pioneering Collection of Open-Source Datasets for Cybersecurity LLM Training

  • 构建覆盖预训练、指令微调、推理蒸馏的全流程开源数据集
  • 持续预训练使综合得分提升15.9%,推理蒸馏助考取CISSP提高15.8%
  • 适合安全研究者和大模型开发者使用,推动领域发展

大语言模型在金融、法律、医疗等领域已取得显著进展,但在网络安全领域仍缺乏高质量的开源预训练语料。为弥补这一缺口,我们推出了覆盖预训练、指令微调和基于自省的推理蒸馏全阶段的综合性数据集。大量消融实验证明其有效性:在公开网络安全基准上,持续预训练可带来15.9%的综合得分提升;推理蒸馏使安全认证(CISSP)成绩提升15.8%。所有数据集与训练好的网络安全大模型将采用ODC-BY和MIT许可证开源,欢迎访问https://huggingface.co/collections/trendmicro-ailab/primus-67b1fd27052b802b4af9d243 获取资源。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have shown remarkable advancements in specialized fields such as finance, law, and medicine. However, in cybersecurity, we have noticed a lack of open-source datasets, with a particular lack of high-quality cybersecurity pretraining corpora, even though much research indicates that LLMs acquire their knowledge during pretraining. To address this, we present a comprehensive suite of datasets covering all major training stages, including pretraining, instruction fine-tuning, and reasoning distillation with cybersecurity-specific self-reflection data. Extensive ablation studies demonstrate their effectiveness on public cybersecurity benchmarks. In particular, continual pre-training on our dataset yields a 15.9% improvement in the aggregate score, while reasoning distillation leads to a 15.8% gain in security certification (CISSP). We will release all datasets and trained cybersecurity LLMs under the ODC-BY and MIT licenses to encourage further research in the community. For access to all datasets and model weights, please refer to https://huggingface.co/collections/trendmicro-ailab/primus-67b1fd27052b802b4af9d243.

网络安全大模型开源数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。