arXiv:2510.27080cs.CRcs.AI2025-10被引 7

用检索增强生成让大模型更快掌握新网络安全威胁。

Adapting Large Language Models to Emerging Cybersecurity using Retrieval Augmented Generation

  • 结合外部数据与检索增强生成,提升模型对新兴威胁的响应能力。
  • 混合检索方法使模型在知识保留和时间推理上表现更优。
  • 适合安全研究人员、漏洞分析师等需要快速应对新威胁的场景。

安全应用正越来越多依赖大语言模型(LLMs)进行网络威胁检测;然而,其推理过程不透明,尤其在需要特定网络安全知识的决策中会降低可信度。由于安全威胁演变迅速,LLMs不仅需回忆历史事件,还需适应新兴漏洞和攻击模式。检索增强生成(RAG)在通用大模型任务中已证明有效,但在网络安全领域的潜力尚未充分探索。本文提出一种基于RAG的框架,旨在通过上下文化网络安全数据,提升LLM在知识保留与时间推理方面的准确性。我们使用外部数据集与Llama-3-8B-Instruct模型,评估基线RAG、优化后的混合检索方法,并在多个性能指标上进行对比分析。结果表明,混合检索在增强LLM对网络安全任务的适应性与可靠性方面具有显著潜力。

原文摘要 · Abstract (English)

Security applications are increasingly relying on large language models (LLMs) for cyber threat detection; however, their opaque reasoning often limits trust, particularly in decisions that require domain-specific cybersecurity knowledge. Because security threats evolve rapidly, LLMs must not only recall historical incidents but also adapt to emerging vulnerabilities and attack patterns. Retrieval-Augmented Generation (RAG) has demonstrated effectiveness in general LLM applications, but its potential for cybersecurity remains underexplored. In this work, we introduce a RAG-based framework designed to contextualize cybersecurity data and enhance LLM accuracy in knowledge retention and temporal reasoning. Using external datasets and the Llama-3-8B-Instruct model, we evaluate baseline RAG, an optimized hybrid retrieval approach, and conduct a comparative analysis across multiple performance metrics. Our findings highlight the promise of hybrid retrieval in strengthening the adaptability and reliability of LLMs for cybersecurity tasks.

大模型网络安全RAG威胁检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。