基于Llama 3.1的8B参数安全专用大模型,提升网络安全任务表现。
Llama-3.1-FoundationAI-SecurityLLM-Base-8B Technical Report
- 在专业网络安全语料上持续预训练,增强安全知识理解能力。
- 在部分安全任务上达到Llama 3.1-70B和GPT-4o-mini水平。
- 适合安全研究人员、企业攻防团队快速部署实用工具。
随着基于Transformer的大语言模型(LLMs)日益融入社会,它们已彻底改变软件工程、创意写作和数字艺术等领域。然而,由于专业训练数据稀缺及网络安全知识表示复杂,其在网络安全领域的应用仍受限。为此,我们提出Foundation-Sec-8B,一个基于Llama 3.1架构、并在精心筛选的网络安全语料上进行持续预训练的专用大模型。我们在多个现有及新设的安全基准上评估该模型,结果显示其在某些网络安全特定任务中表现可媲美Llama 3.1-70B和GPT-4o-mini。通过公开发布该模型,我们旨在加速人工智能驱动工具在公共与私有网络安全场景中的发展与应用。
原文摘要 · Abstract (English)
As transformer-based large language models (LLMs) increasingly permeate society, they have revolutionized domains such as software engineering, creative writing, and digital arts. However, their adoption in cybersecurity remains limited due to challenges like scarcity of specialized training data and complexity of representing cybersecurity-specific knowledge. To address these gaps, we present Foundation-Sec-8B, a cybersecurity-focused LLM built on the Llama 3.1 architecture and enhanced through continued pretraining on a carefully curated cybersecurity corpus. We evaluate Foundation-Sec-8B across both established and new cybersecurity benchmarks, showing that it matches Llama 3.1-70B and GPT-4o-mini in certain cybersecurity-specific tasks. By releasing our model to the public, we aim to accelerate progress and adoption of AI-driven tools in both public and private cybersecurity contexts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。