构建首个覆盖81种攻击技术的系统日志数据集,助力大模型自动解析安全告警。
CAM-LDS: Cyber Attack Manifestations for Automatic Interpretation of System Logs and Security Alerts
- 构建涵盖13类战术、81种攻击技术的开源日志数据集CAM-LDS。
- 使用大模型分析发现约1/3攻击步骤可被完全正确识别,另1/3识别充分。
- 适合研究日志自动化分析、大模型在网络安全中应用的团队使用。
系统日志对入侵检测和溯源至关重要,但人工分析面临数据量大、格式异构、信息非结构化等挑战。现有自动化方法多依赖领域定制配置,如专家定义规则、手工日志解析器或人工特征工程,难以实现语义理解与因果解释。相比之下,大语言模型(LLM)具备跨领域、跨格式的解释能力。然而,该方向受限于缺乏覆盖广泛攻击技术的公开标注数据集。为此,本文提出网络攻击表现日志数据集(CAM-LDS),包含7个攻击场景,覆盖13个战术、81种具体技术,数据源自18个不同来源,在全开源可复现环境中采集。我们提取攻击执行直接产生的日志事件,用于分析命令可观测性、事件频率、性能指标及入侵检测告警。进一步通过案例研究展示使用LLM处理CAM-LDS的效果:约三分之一攻击步骤被完全正确预测,另三分之一得到充分识别,验证了基于大模型的日志解释潜力及本数据集的实用价值。
原文摘要 · Abstract (English)
Log data are essential for intrusion detection and forensic investigations. However, manual log analysis is tedious due to high data volumes, heterogeneous event formats, and unstructured messages. Even though many automated methods for log analysis exist, they usually still rely on domain-specific configurations such as expert-defined detection rules, handcrafted log parsers, or manual feature-engineering. Crucially, the level of automation of conventional methods is limited due to their inability to semantically understand logs and explain their underlying causes. In contrast, Large Language Models enable domain- and format-agnostic interpretation of system logs and security alerts. Unfortunately, research on this topic remains challenging, because publicly available and labeled data sets covering a broad range of attack techniques are scarce. To address this gap, we introduce the Cyber Attack Manifestation Log Data Set (CAM-LDS), comprising seven attack scenarios that cover 81 distinct techniques across 13 tactics and collected from 18 distinct sources within a fully open-source and reproducible test environment. We extract log events that directly result from attack executions to facilitate analysis of manifestations concerning command observability, event frequencies, performance metrics, and intrusion detection alerts. We further present an illustrative case study utilizing an LLM to process the CAM-LDS. The results indicate that correct attack techniques are predicted perfectly for approximately one third of attack steps and adequately for another third, highlighting the potential of LLM-based log interpretation and utility of our data set.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。