arXiv:2506.12100cs.CRcs.AI2025-06被引 3

用嵌入向量分析大模型生成内容中内部知识与外部检索的贡献比例。

LLM Embedding-based Attribution (LEA): Quantifying Source Contributions to Generative Model's Response for Vulnerability Analysis

  • 通过嵌入相似度量化模型输出中内部知识与检索内容的贡献度。
  • 在500个漏洞上验证,大模型对有效检索的依赖度超过95%准确识别。
  • 帮助安全分析师审计AI决策可信度,防范错误检索带来的风险。

大型语言模型(LLMs)在网络安全威胁分析中的应用日益广泛,但其在高安全敏感环境中的部署引发信任与安全担忧。2025年已披露超21,000个漏洞,人工分析已不可行,亟需可扩展且可验证的AI支持。当查询模型时,新兴漏洞因训练数据存在截止日期而难以应对。尽管检索增强生成(RAG)可注入最新上下文以缓解此问题,但尚不清楚模型对检索证据与内部知识的依赖程度,以及检索内容是否真实有效。这种不确定性可能误导安全分析师,误判补丁优先级,增加安全风险。为此,本文提出基于嵌入的归因方法(LEA),用于分析生成响应中漏洞利用分析的来源贡献。具体而言,LEA量化了内部知识与检索内容在生成结果中的相对贡献。我们在2016至2025年间披露的500个关键漏洞上,对三种RAG设置(有效、通用、错误)及三种前沿大模型进行了评估。结果表明,LEA在大模型上对非检索、通用检索和有效检索场景的区分准确率超过95%。最后,我们揭示了错误检索带来的局限性,并警示网络安全社区避免盲目依赖大模型与RAG进行漏洞分析。LEA为安全分析师提供了审计RAG增强工作流的指标,提升人工智能在网络安全威胁分析中透明与可信的部署水平。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are increasingly used for cybersecurity threat analysis, but their deployment in security-sensitive environments raises trust and safety concerns. With over 21,000 vulnerabilities disclosed in 2025, manual analysis is infeasible, making scalable and verifiable AI support critical. When querying LLMs, dealing with emerging vulnerabilities is challenging as they have a training cut-off date. While Retrieval-Augmented Generation (RAG) can inject up-to-date context to alleviate the cut-off date limitation, it remains unclear how much LLMs rely on retrieved evidence versus the model's internal knowledge, and whether the retrieved information is meaningful or even correct. This uncertainty could mislead security analysts, mis-prioritize patches, and increase security risks. Therefore, this work proposes LLM Embedding-based Attribution (LEA) to analyze the generated responses for vulnerability exploitation analysis. More specifically, LEA quantifies the relative contribution of internal knowledge vs. retrieved content in the generated responses. We evaluate LEA on 500 critical vulnerabilities disclosed between 2016 and 2025, across three RAG settings -- valid, generic, and incorrect -- using three state-of-the-art LLMs. Our results demonstrate LEA's ability to detect clear distinctions between non-retrieval, generic-retrieval, and valid-retrieval scenarios with over 95% accuracy on larger models. Finally, we demonstrate the limitations posed by incorrect retrieval of vulnerability information and raise a cautionary note to the cybersecurity community regarding the blind reliance on LLMs and RAG for vulnerability analysis. LEA offers security analysts with a metric to audit RAG-enhanced workflows, improving the transparent and trustworthy deployment of AI in cybersecurity threat analysis.

大模型安全漏洞分析RAG审计可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。