arXiv:2509.00081cs.CRcs.AI2025-09

用大模型+领域知识图谱提升网络安全日志的精准解析与可解释性。

Enabling Transparent Cyber Threat Intelligence Combining Large Language Models and Domain Ontologies

  • 结合领域本体与SHACL约束,引导大模型输出结构化语义数据。
  • 在公开数据集上信息提取准确率显著高于传统提示方法。
  • 适合安全分析师、威胁情报系统开发者使用,提升分析可信度。

有效的网络安全威胁情报依赖于从安全系统日志中提取出准确结构化且语义丰富的信息。然而,现有方法在识别和解释恶意事件时往往难以做到可靠且透明,尤其面对非结构化或模糊的日志条目。本文提出一种新方法,将本体驱动的结构化输出与大语言模型(LLM)相结合,构建一个能提升日志信息提取准确率与可解释性的AI代理。核心在于利用领域本体与基于SHACL的约束,指导语言模型输出结构并确保生成图谱的语义有效性。提取的信息被组织为本体增强的图数据库,支持后续语义分析与查询。该方法设计基于蜜罐日志数据的分析需求,其以大量恶意活动为特征。尽管案例研究聚焦于此场景,实验评估采用公开可用数据集。结果表明,相较于传统仅靠提示的方法,本方法在信息提取准确性上表现更优,重点强调提取质量而非处理速度。

原文摘要 · Abstract (English)

Effective Cyber Threat Intelligence (CTI) relies upon accurately structured and semantically enriched information extracted from cybersecurity system logs. However, current methodologies often struggle to identify and interpret malicious events reliably and transparently, particularly in cases involving unstructured or ambiguous log entries. In this work, we propose a novel methodology that combines ontology-driven structured outputs with Large Language Models (LLMs), to build an Artificial Intelligence (AI) agent that improves the accuracy and explainability of information extraction from cybersecurity logs. Central to our approach is the integration of domain ontologies and SHACL-based constraints to guide the language model's output structure and enforce semantic validity over the resulting graph. Extracted information is organized into an ontology-enriched graph database, enabling future semantic analysis and querying. The design of our methodology is motivated by the analytical requirements associated with honeypot log data, which typically comprises predominantly malicious activity. While our case study illustrates the relevance of this scenario, the experimental evaluation is conducted using publicly available datasets. Results demonstrate that our method achieves higher accuracy in information extraction compared to traditional prompt-only approaches, with a deliberate focus on extraction quality rather than processing speed.

威胁情报大模型知识图谱安全分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。