六智能体协作自动评估企业安全风险,15分钟完成,准确率达85%。
An Agentic Multi-Agent Architecture for Cybersecurity Risk Management
- 六个智能体分阶段协作,共享上下文逐步推进风险分析。
- 在医疗企业测试中与专家判断一致率达85%,覆盖92%风险点。
- 领域微调模型能识别特定行业风险,但上下文容量限制了部署。
为小企业获取真实网络安全风险评估成本高昂——符合NIST CSF标准的评估至少需1.5万美元、耗时数周,且依赖稀缺的专业人员,多数企业因此放弃。我们构建了一个六智能体系统,分别负责组织画像、资产映射、威胁分析、控制评估、风险评分和建议生成。各智能体共享持久化上下文,后序阶段基于前期结论,区别于传统顺序式智能体流水线。我们在一家15人的合规医疗公司上测试,与三位CISSP专家独立评估结果对比,系统在严重性分类上达成85%一致性,覆盖92%已识别风险,全程用时不足15分钟。随后在五个模拟但行业真实的组织(医疗、金融科技、制造、零售、SaaS)上运行30次单智能体评估,对比通用模型Mistral-7B与领域微调模型。两者均成功完成所有任务;微调模型识别出基线模型完全未发现的风险:医疗中的PHI泄露、制造中的OT/IIoT漏洞、零售平台特有风险。然而,完整多智能体流水线在配备4096令牌默认上下文窗口的Tesla T4上30次尝试全部失败,表明上下文容量是关键瓶颈。
原文摘要 · Abstract (English)
Getting a real cybersecurity risk assessment for a small organization is expensive -- a NIST CSF-aligned engagement runs $15,000 on the low end, takes weeks, and depends on practitioners who are genuinely scarce. Most small companies skip it entirely. We built a six-agent AI system where each agent handles one analytical stage: profiling the organization, mapping assets, analyzing threats, evaluating controls, scoring risks, and generating recommendations. Agents share a persistent context that grows as the assessment proceeds, so later agents build on what earlier ones concluded -- the mechanism that distinguishes this from standard sequential agent pipelines. We tested it on a 15-person HIPAA-covered healthcare company and compared outputs to independent assessments by three CISSP practitioners -- the system agreed with them 85% of the time on severity classifications, covered 92% of identified risks, and finished in under 15 minutes. We then ran 30 repeated single-agent assessments across five synthetic but sector-realistic organizational profiles in healthcare, fintech, manufacturing, retail, and SaaS, comparing a general-purpose Mistral-7B against a domain fine-tuned model. Both completed every run. The fine-tuned model flagged threats the baseline could not see at all: PHI exposure in healthcare, OT/IIoT vulnerabilities in manufacturing, platform-specific risks in retail. The full multi-agent pipeline, however, failed every one of 30 attempts on a Tesla T4 with its 4,096-token default context window -- context capacity, not model quality, turned out to be the binding constraint.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。