用大模型自动构建带功能属性的溯源图,提升攻击检测与解释能力
An End-to-End Framework for Functionality-Embedded Provenance Graph Construction and Threat Interpretation
- 基于大模型自动生成规则,从异构日志中构建溯源图
- 为系统实体注入功能上下文,使检测器性能平均提升18%
- 生成自然语言攻击摘要,支持分析师快速研判
溯源图通过日志建模系统级因果关系,使异常检测器学习正常行为并识别偏离。然而现有方法依赖脆弱的手工规则构建溯源图,缺乏系统实体的功能上下文,且难以支持分析员调查。我们提出Auto-Prov,一个自适应的端到端框架,利用大语言模型(LLMs)从异构且动态演化的日志中自动构建溯源图,嵌入系统级功能属性,使基于溯源图的异常检测器能从中学习,并将检测到的攻击总结为自然语言文本以辅助分析。Auto-Prov通过自动生成规则聚类未见日志类型,高效提取溯源边和实体信息;进一步结合大模型推理与行为估计,推断已知及未知系统实体的系统级功能上下文。在四种先进溯源图基检测器上评估显示,Auto-Prov持续提升检测性能,跨异构日志格式具有良好泛化性,生成稳定可解释的攻击摘要,在系统演化下仍保持鲁棒性。
原文摘要 · Abstract (English)
Provenance graphs model causal system-level interactions from logs, enabling anomaly detectors to learn normal behavior and detect deviations as attacks. However, existing approaches rely on brittle, manually engineered rules to build provenance graphs, lack functional context for system entities, and provide limited support for analyst investigation. We present Auto-Prov, an adaptive, end-to-end framework that leverages large language models (LLMs) to automatically construct provenance graphs from heterogeneous and evolving logs, embed system-level functional attributes into the graph, enable provenance graph-based anomaly detectors to learn from these enriched graphs, and summarize the detected attacks to assist an analyst's investigation. Auto-Prov clusters unseen log types and efficiently extracts provenance edges and entity-level information via automatically generated rules. It further infers system-level functional context for both known and previously unseen system entities using a combination of LLM inference and behavior-based estimation. Attacks detected by provenance-graph-based anomaly detectors trained on Auto-Prov's graphs are then summarized into natural-language text. We evaluate Auto-Prov with four state-of-the-art provenance graph-based detectors across diverse logs. Results show that Auto-Prov consistently enhances detection performance, generalizes across heterogeneous log formats, and produces stable, interpretable attack summaries that remain robust under system evolution.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。