用多智能体大模型自动生成逼真内鬼行为数据,解决检测训练数据匮乏难题。
Chimera: Harnessing Multi-Agent LLMs for Automatic Insider Threat Simulation
- 构建多智能体系统模拟员工角色与互动,生成真实组织动态。
- 基于15起真实事件生成新数据集ChimeraLog,更难检测且具多样性。
- 适合安全研究者和内鬼检测算法开发者使用,推动真实场景评估。
内部威胁是企业环境中持续且严重的安全风险,但因其恶意行为常隐藏于正常操作中而难以发现。尽管基于机器学习的内鬼检测(ITD)方法已显成效,其性能受限于高质量、真实行为数据的稀缺。企业内部数据敏感难以获取,现有公开或合成数据集规模小、缺乏真实性、语义丰富度和行为多样性。为此,我们提出Chimera——一种基于大语言模型的多智能体框架,可自动模拟良性与恶意内鬼行为,并在多种企业场景下生成完整系统日志。每个智能体代表具有细粒度角色的员工,支持团队会议、成对交互与自主调度,以捕捉真实的组织动态。基于15起真实事件抽象出的内鬼攻击,我们在三种典型数据敏感组织场景中部署Chimera,构建了ChimeraLog新数据集,用于开发与评估ITD方法。通过人类评估与定量分析验证其多样性和真实性。实验表明,现有ITD方法在ChimeraLog上的检测性能显著低于以往数据集,说明其更具挑战性与现实性。此外,即便存在分布偏移,基于ChimeraLog训练的模型仍表现出强泛化能力,凸显大语言模型多智能体仿真对推进内鬼检测的实际价值。
原文摘要 · Abstract (English)
Insider threats pose a persistent and critical security risk, yet are notoriously difficult to detect in complex enterprise environments, where malicious actions are often hidden within seemingly benign user behaviors. Although machine-learning-based insider threat detection (ITD) methods have shown promise, their effectiveness is fundamentally limited by the scarcity of high-quality and realistic training data. Enterprise internal data is highly sensitive and rarely accessible, while existing public and synthetic datasets are either small-scale or lack sufficient realism, semantic richness, and behavioral diversity. To address this challenge, we propose Chimera, an LLM-based multi-agent framework that automatically simulates both benign and malicious insider activities and generates comprehensive system logs across diverse enterprise environments. Chimera models each agent as an individual employee with fine-grained roles and supports group meetings, pairwise interactions, and self-organized scheduling to capture realistic organizational dynamics. Based on 15 insider attacks abstracted from real-world incidents, we deploy Chimera in three representative data-sensitive organizational scenarios and construct ChimeraLog, a new dataset for developing and evaluating ITD methods. We evaluate ChimeraLog through human studies and quantitative analyses, demonstrating its diversity and realism. Experiments with existing ITD methods show substantially lower detection performance on ChimeraLog compared to prior datasets, indicating a more challenging and realistic benchmark. Moreover, despite distribution shifts, models trained on ChimeraLog exhibit strong generalization, highlighting the practical value of LLM-based multi-agent simulation for advancing insider threat detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。