用可解释性分析让大模型读懂网络攻击行为,提升安全检测可信度。
Attribution-Driven Explainable Intrusion Detection with Encoder-Based Large Language Models

- 基于注意力机制分析大模型决策依据,识别关键流量模式。
- 发现模型判断依赖真实攻击特征,与传统检测原则一致。
- 适合关注AI安全可信性的研究人员与网络安全工程师。
软件定义网络(SDN)提升了网络灵活性,但也增加了对可靠且可解释入侵检测的需求。尽管大型语言模型(LLMs)凭借强大的表征学习能力被用于网络安全任务,但其缺乏透明性限制了在安全关键场景中的应用。理解LLM的决策过程至关重要。本文针对基于编码器的LLMs,利用流级流量特征开展归因驱动的入侵检测分析。归因分析表明,模型决策由有意义的流量行为模式驱动,提升了基于Transformer的SDN入侵检测的透明度与可信度。这些模式与既有的入侵检测原则相符,说明LLMs能从流量动态中学习攻击行为。本工作展示了归因方法在验证和信任基于LLM的安全分析中的价值。
原文摘要 · Abstract (English)
Software-Defined Networking (SDN) improves network flexibility but also increases the need for reliable and interpretable intrusion detection. Large Language Models (LLMs) have recently been explored for cybersecurity tasks due to their strong representation learning capabilities; however, their lack of transparency limits their practical adoption in security-critical environments. Understanding how LLMs make decisions is therefore essential. This paper presents an attribution-driven analysis of encoder-based LLMs for network intrusion detection using flow-level traffic features. Attribution analysis demonstrates that model decisions are driven by meaningful traffic behavior patterns, improving transparency and trust in transformer-based SDN intrusion detection. These patterns align with established intrusion detection principles, indicating that LLMs learn attack behavior from traffic dynamics. This work demonstrates the value of attribution methods for validating and trusting LLM-based security analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。