用多智能体模拟与心理推理提升企业内鬼检测准确率
A Hybrid Insider Threat Detection Framework Combining Multi-Agent Simulation, Layered SIEM Correlation, and Theory-of-Mind Reasoning
- 融合多智能体仿真与心智理论,从邮件行为推断异常意图
- 检测准确率最高达0.944,误报减少至每轮0.0条
- 适合安全团队部署,尤其关注低误报高可信的场景
本文提出一种混合内鬼检测框架,整合多智能体仿真、分层SIEM关联、信任自适应阈值、行为与通信取证及心智理论推理。将邮件视为协调与社会工程证据流,关联认证、文件访问和权限事件。评估四种变体:LSC、CE-SIEM、EG-SIEM 和 EG-SIEM-Enron(基于Enron校准的邮件取证)。在十轮匹配实验中,八名恶意用户检测的个体级F1得分从LSC的0.567提升至CE-SIEM的0.774、EG-SIEM的0.898,最终达EG-SIEM-Enron的0.944;经霍姆-邦弗朗尼校正的配对威尔科克斯检验确认前三者差异显著。证据门控使确认误报从LSC的33.7条/轮、CE-SIEM的49.3条/轮降至0.2条和0.0条,精度提升但确认时间延长。域迁移测试显示,基于Enron训练的邮件分类器无法跨域迁移,但目标域微调后在分组模板族评估下达到F1=0.974。在CERT r4.2数据集上,证据门控逻辑提升个体检测性能并降低误报,优于表格异常检测器与句子嵌入基线。扩展至1000个代理的可扩展性测试表明检测质量稳定。
原文摘要 · Abstract (English)
This paper presents a hybrid insider threat detection framework for enterprise environments, integrating multi-agent simulation, layered SIEM correlation, trust-adaptive thresholds, behavioral and communication forensics, and Theory-of-Mind reasoning. Email is treated not as a control channel but as a coordination and social-engineering evidence stream correlated with authentication, file-access, and privilege events. Four variants are evaluated: Layered SIEM-Core (LSC), Cognitive-Enriched SIEM (CE-SIEM), Evidence-Gated SIEM (EG-SIEM), and EG-SIEM-Enron with Enron-calibrated email forensics. Across ten matched runs with eight malicious insiders, actor-level F1 improves from 0.567 for LSC to 0.774 for CE-SIEM, 0.898 for EG-SIEM, and 0.944 for EG-SIEM-Enron; paired Wilcoxon tests confirm the first three differences after Holm-Bonferroni correction. Evidence gating reduces confirmed false positives from 33.7 per run under LSC and 49.3 under CE-SIEM to 0.2 and 0.0 respectively, a precision gain traded against longer confirmation time. Domain-shift evaluation shows that an Enron-trained email classifier does not transfer to a different operational email domain, although target-domain fine-tuning reaches F1 = 0.974 under grouped template-family evaluation. On CERT r4.2, the evidence-gated logic improves actor-level detection while reducing false positives, outperforming tabular anomaly detectors and a sentence-embedding baseline. Scalability tests to 1000 agents indicate stable detection quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。