用多智能体框架提升网络安全问答的准确性和可解释性。
MITRE-SAGE: A Multi-Agent Cybersecurity Question-Answering Model

- 分三步处理:理解问题、检索证据、合成答案,增强推理能力。
- 在8个任务中5项超越基线,轻量版模型达最优表现。
- 适合安全分析、漏洞评估等场景,为研究提供标准评测集。
有效网络安全运营需要及时准确地分析大规模异构安全信息;然而,分析师正面临信息过载、告警疲劳和时间受限决策的挑战。尽管大语言模型(LLMs)在问答任务中展现出潜力,其在网络安全领域的应用仍受限于领域知识不足、幻觉倾向以及难以捕捉语义与结构关系。本文提出MITRE-SAGE,一种融合语义与结构化网络安全知识的多智能体检索增强生成框架,以提升基于LLM的问答系统可靠性与可解释性。通过将复杂任务分解为查询理解、证据检索与答案合成,MITRE-SAGE有效支持漏洞评估、威胁画像和关系提取等任务。此外,我们构建了MITRE-QA,一个包含3,000个问答对的综合性基准,用于系统评估不同方法在多样化网络安全任务上的表现。大量实验表明,MITRE-SAGE持续优于独立的LLM和传统RAG方法。值得注意的是,由Qwen2.5-7B子智能体与Qwen2.5-14B协调器组成的轻量配置在8个任务中的5项上取得最佳性能,验证了该多智能体框架的有效性。结果表明,MITRE-SAGE是一种可扩展且可解释的可靠网络安全问答方案,而MITRE-QA为未来研究提供了标准化评估基准。
原文摘要 · Abstract (English)
Effective cybersecurity operations require timely and accurate analysis of large-scale heterogeneous security information; however, analysts increasingly struggle with information overload, alert fatigue, and time-constrained decision-making. Although large language models (LLMs) have demonstrated promising capabilities for question answering (QA), their effectiveness in cybersecurity remains limited by insufficient domain knowledge, a tendency to hallucinate, and difficulties in capturing both semantic and structural relationships. This work proposes MITRE-SAGE, a multi-agent retrieval-augmented generation framework that integrates semantic and structural cybersecurity knowledge to improve the reliability and interpretability of LLM-based QA systems. By decomposing complex tasks into query interpretation, evidence retrieval, and answer synthesis, MITRE-SAGE effectively supports cybersecurity tasks such as vulnerability assessment, threat profiling, and relationship extraction. Furthermore, we propose MITRE-QA, a comprehensive benchmark comprising 3,000 question-answer pairs for evaluating LLMs across diverse cybersecurity knowledge tasks, and use it to systematically evaluate MITRE-SAGE against representative baseline methods. Extensive experiments demonstrate that MITRE-SAGE consistently outperforms standalone LLMs and conventional RAG approaches. Notably, a lightweight configuration comprising Qwen2.5-7B sub-agents and a Qwen2.5-14B orchestrator achieves superior performance on five of the eight benchmark tasks, indicating the effectiveness of the proposed multi-agent framework. The results highlight the potential of MITRE-SAGE as a scalable and interpretable approach for reliable cybersecurity QA, while MITRE-QA provides a standardized benchmark for future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。