arXiv:2604.27321cs.CRcs.AI2026-04被引 2

用大模型自动处理安全事件,提速降错,三步搞定威胁检测、查询生成和处置。

Toward Autonomous SOC Operations: End-to-End LLM Framework for Threat Detection, Query Generation, and Resolution in Security Operations

论文配图:Toward Autonomous SOC Operations: End-to-End LLM Framework for Threat Detection, Query Generation, and Resolution in Security Operations
图 1 · 摘自论文原文
  • 集成多模型检测与语法约束查询生成,提升安全日志分析准确性。
  • 将事件处置时间从数小时缩短至10分钟内,准确率提升至90%。
  • 适合需要自动化响应的网络安全团队,尤其适配IBM QRadar等平台。

安全运营中心(SOC)面临威胁量激增、平台异构及人工研判耗时等挑战。本文提出端到端威胁管理框架,融合集成检测、语法约束查询生成与检索增强的处置支持。检测模块综合三个最优大语言模型(LLM),在SIEM日志上实现82.8%准确率与0.120假阳性率。提出SQM(语法查询元数据)架构,基于平台语法约束、元数据检索与文档引导提示,生成适用于IBM QRadar与Google SecOps的可执行查询,BLEU得分为0.384,ROUGE-L为0.731,性能超过基线两倍以上。结合SQM证据后,处置代码预测准确率由78.3%提升至90.0%,推荐质量评分达8.70。在真实生产环境中,平均事件研判时间从数小时降至10分钟以内。结果表明,领域受限且带检索增强的大模型架构,可满足大规模安全运营对可靠性和效率的严苛要求。

原文摘要 · Abstract (English)

Security Operations Centers (SOCs) face mounting operational challenges. These challenges come from increasing threat volumes, heterogeneous SIEM platforms, and time-consuming manual triage workflows. We present an end-to-end threat management framework that integrates ensemble-based detection, syntax-constrained query generation, and retrieval-augmented resolution support to automate critical security workflows. Our detection module evaluates both traditional machine learning classifiers and large language models (LLMs), then combines the three best-performing LLMs to create an ensemble model, achieving 82.8% accuracy while maintaining 0.120 false positive rate on SIEM logs. We introduce the SQM (Syntax Query Metadata) architecture for automated evidence collection. It uses platform-specific syntax constraints, metadata-based retrieval, and documentation-grounded prompting to generate executable queries for IBM QRadar and Google SecOps. SQM achieves a BLEU score of 0.384 and a ROUGE-L score of 0.731. These results are more than twice as good as the baseline LLM performance. For incident resolution and recommendation generation, we demonstrate that integrating SQM-derived evidence improves resolution code prediction accuracy from 78.3% to 90.0%, with an overall recommendation quality score of 8.70. In production SOC environments, our framework reduces average incident triage time from hours to under 10 minutes. This work demonstrates that domain-constrained LLM architectures with retrieval augmentation can meet the strict reliability and efficiency requirements of operational security environments at scale.

安全自动化大模型应用SOC优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。