用联邦对抗训练解决多地点内鬼检测的数据偏斜问题。
FedAT: Federated Adversarial Training for Distributed Insider Threat Detection
- 联邦学习结合生成模型缓解跨客户端数据分布不均问题。
- 新方法在多类内鬼检测任务中显著提升准确率。
- 适合关注隐私保护与分布式安全的机构应用。
内鬼威胁通常源自组织内部,攻击者是与组织密切关联的实体。通过分析其对权限资源的操作序列,可识别内鬼行为。近年来,基于机器学习的内鬼检测(ITD)受到广泛关注。然而,多数方法采用集中式建模,难以适用于多地点运营的组织,因用户行为数据涉及隐私无法共享。此外,各位置分散的数据导致极端类别不平衡。联邦学习(FL)作为一种分布式建模范式,近年来备受关注,但其在实际场景中的内鬼检测应用仍待研究。本文提出一种支持非独立同分布(non-IID)数据的联邦多类内鬼检测框架,设计了联邦对抗训练(FedAT)方法,利用生成模型缓解客户端间数据偏斜问题;同时引入自归一化神经网络多层感知机(SNN-MLP)模型以提升检测性能。通过全面实验与基准对比,验证了所提方案在准确率和鲁棒性上的优势。
原文摘要 · Abstract (English)
Insider threats usually occur from within the workplace, where the attacker is an entity closely associated with the organization. The sequence of actions the entities take on the resources to which they have access rights allows us to identify the insiders. Insider Threat Detection (ITD) using Machine Learning (ML)-based approaches gained attention in the last few years. However, most techniques employed centralized ML methods to perform such an ITD. Organizations operating from multiple locations cannot contribute to the centralized models as the data is generated from various locations. In particular, the user behavior data, which is the primary source of ITD, cannot be shared among the locations due to privacy concerns. Additionally, the data distributed across various locations result in extreme class imbalance due to the rarity of attacks. Federated Learning (FL), a distributed data modeling paradigm, gained much interest recently. However, FL-enabled ITD is not yet explored, and it still needs research to study the significant issues of its implementation in practical settings. As such, our work investigates an FL-enabled multiclass ITD paradigm that considers non-Independent and Identically Distributed (non-IID) data distribution to detect insider threats from different locations (clients) of an organization. Specifically, we propose a Federated Adversarial Training (FedAT) approach using a generative model to alleviate the extreme data skewness arising from the non-IID data distribution among the clients. Besides, we propose to utilize a Self-normalized Neural Network-based Multi-Layer Perceptron (SNN-MLP) model to improve ITD. We perform comprehensive experiments and compare the results with the benchmarks to manifest the enhanced performance of the proposed FedATdriven ITD scheme.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。