arXiv:2509.23573cs.CRcs.AI2025-09被引 5

揭示大模型在网络安全情报中因威胁环境复杂导致的三大认知缺陷

Uncovering Vulnerabilities of LLM-Assisted Cyber Threat Intelligence

  • 构建人机协同分类框架,精准识别网络安全情报中的失败模式
  • 发现三类特定缺陷:虚假关联、信息冲突、对新威胁泛化能力差
  • 通过因果干预验证防御策略,为安全智能体设计提供可落地方案

大型语言模型(LLMs)正被广泛用于协助安全分析师应对激增的网络威胁,自动化完成漏洞评估到事件响应等任务。然而在实际网络安全情报(CTI)工作流中,可靠性问题依然显著。现有解释多归因于通用模型缺陷(如幻觉),我们提出主要瓶颈在于威胁环境本身的异构性、波动性和碎片化。在此背景下,证据具有交织性、众包性和时间不稳定性,而这些特性极少被现有基于LLM的研究捕捉。本文开展了一项全面的实证研究,揭示了LLM在CTI推理中的脆弱性。我们提出一种人机协同的分类框架,可靠标注整个CTI生命周期中的故障模式,避免了自动化‘以LLM为裁判’流程的脆弱性。识别出三类领域特异性认知失败:由表面元数据引发的虚假相关、来自冲突来源的矛盾知识、对新兴威胁的泛化受限。通过因果干预验证这些机制,并证明针对性防御可显著降低失败率。这些结果共同为构建鲁棒、领域感知的CTI智能体提供了具体路线图。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly used to help security analysts manage the surge of cyber threats, automating tasks from vulnerability assessment to incident response. Yet in operational CTI workflows, reliability gaps remain substantial. Existing explanations often point to generic model issues (e.g., hallucination), but we argue the dominant bottleneck is the threat landscape itself: CTI is heterogeneous, volatile, and fragmented. Under these conditions, evidence is intertwined, crowdsourced, and temporally unstable, which are properties that standard LLM-based studies rarely capture. In this paper, we present a comprehensive empirical study of LLM vulnerabilities in CTI reasoning. We introduce a human-in-the-loop categorization framework that robustly labels failure modes across the CTI lifecycle, avoiding the brittleness of automated "LLM-as-a-judge" pipelines. We identify three domain-specific cognitive failures: spurious correlations from superficial metadata, contradictory knowledge from conflicting sources, and constrained generalization to emerging threats. We validate these mechanisms via causal interventions and show that targeted defenses reduce failure rates significantly. Together, these results offer a concrete roadmap for building resilient, domain-aware CTI agents.

大模型安全网络安全认知缺陷智能情报

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。