arXiv:2510.00240cs.CRcs.AI2025-10被引 9

专为网络安全设计的高效语言模型,提升威胁分析与漏洞检测精度。

SecureBERT 2.0: Advanced Language Model for Cybersecurity Intelligence

  • 基于ModernBERT架构,增强长文本与代码混合文档的编码能力。
  • 在超130亿文本+5300万代码令牌上预训练,性能超越现有基准。
  • 适合安全分析师、自动化威胁检测系统开发者使用。

高效分析网络安全与威胁情报数据需要能够理解专业术语、复杂文档结构及自然语言与源代码相互依赖的语言模型。编码器仅用的Transformer架构能提供高效稳健的表示,支持语义搜索、技术实体抽取和语义分析等关键任务,对自动化威胁检测、事件分类和漏洞评估至关重要。然而,通用语言模型常缺乏高精度所需的领域特定适应性。我们提出SecureBERT 2.0,一种专为网络安全应用优化的增强型编码器仅用语言模型。基于ModernBERT架构,SecureBERT 2.0引入改进的长上下文建模与分层编码,有效处理扩展且异构的文档,包括威胁报告和源代码片段。在比前代大十三倍以上的领域专用语料库上预训练,包含超过130亿文本令牌和5300万代码令牌,来自多样真实来源。实验表明,SecureBERT 2.0在多个网络安全基准测试中达到最先进水平,在威胁情报语义搜索、语义分析、网络安全命名实体识别以及代码中自动漏洞检测方面均有显著提升。

原文摘要 · Abstract (English)

Effective analysis of cybersecurity and threat intelligence data demands language models that can interpret specialized terminology, complex document structures, and the interdependence of natural language and source code. Encoder-only transformer architectures provide efficient and robust representations that support critical tasks such as semantic search, technical entity extraction, and semantic analysis, which are key to automated threat detection, incident triage, and vulnerability assessment. However, general-purpose language models often lack the domain-specific adaptation required for high precision. We present SecureBERT 2.0, an enhanced encoder-only language model purpose-built for cybersecurity applications. Leveraging the ModernBERT architecture, SecureBERT 2.0 introduces improved long-context modeling and hierarchical encoding, enabling effective processing of extended and heterogeneous documents, including threat reports and source code artifacts. Pretrained on a domain-specific corpus more than thirteen times larger than its predecessor, comprising over 13 billion text tokens and 53 million code tokens from diverse real-world sources, SecureBERT 2.0 achieves state-of-the-art performance on multiple cybersecurity benchmarks. Experimental results demonstrate substantial improvements in semantic search for threat intelligence, semantic analysis, cybersecurity-specific named entity recognition, and automated vulnerability detection in code within the cybersecurity domain.

网络安全语言模型代码理解实体识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。