用生成式AI和联邦学习提升入侵检测系统隐私与泛化能力
Generative AI and Federated Learning for Intrusion Detection Systems: A Survey
- 结合生成模型与联邦学习,实现分布式隐私保护下的异常检测
- 生成数据可缓解样本不平衡与真实数据难获取问题
- 适合关注网络安全、隐私计算及智能检测的研究者
入侵检测系统(IDS)在现代网络物理系统、物联网、企业及分布式网络环境中对监控流量、识别恶意行为至关重要。但构建可靠的IDS模型仍面临诸多挑战:攻击行为持续演化,真实数据集难以获取,流量记录不完整,攻击类别严重不平衡,且隐私限制导致无法集中收集数据。近年来,生成式人工智能(Generative AI)与联邦学习(FL)的发展为解决这些问题提供了新思路。生成模型可用于异常检测、合成流量生成、数据增强、缺失数据补全、对抗性流量生成以及检测告警解释;联邦学习则可在不直接共享本地网络流量的前提下实现分布式训练,适用于隐私敏感和地理分布式的场景。本文系统综述了生成式AI与联邦学习在IDS中的应用,梳理了代表性研究方向,包括对抗机器学习、基于异常的检测、面向物联网的IDS、可解释的IDS及基准数据集。进一步按模型家族与任务目标分类讨论生成式AI的应用,涵盖自编码器、生成对抗网络(GANs)、扩散模型与大语言模型(LLMs)。最后,综述了生成式AI与联邦学习融合的新兴研究,探讨了合成数据质量、真实流量生成、双重用途对抗风险、非独立同分布客户端分布、通信高效模型共享、联邦IDS基准测试以及领域专用大语言模型在网络安全中的应用等开放挑战。
原文摘要 · Abstract (English)
Intrusion Detection Systems (IDSs) are essential for monitoring network traffic and identifying malicious activities in modern cyber-physical, Internet of Things (IoT), enterprise, and distributed network environments. However, developing reliable IDS models remains challenging because attack behaviors evolve over time, realistic datasets are difficult to obtain, traffic records may be incomplete, attack classes are often imbalanced, and privacy constraints limit centralized data collection. Recent advances in generative artificial intelligence (AI) and Federated Learning (FL) provide new opportunities to address these limitations. Generative models can support anomaly detection, synthetic traffic generation, data augmentation, data imputation, adversarial traffic generation, and IDS alert explanation. FL enables distributed IDS training without directly sharing local network traffic, making it suitable for privacy-sensitive and geographically distributed environments. This survey provides a structured review of generative AI and FL techniques for IDS. We first summarize representative IDS research directions, including adversarial machine learning, anomaly-based detection, IoT-oriented IDS, explainable IDS, and benchmark datasets. We then categorize generative AI applications in IDS according to model families and task objectives, covering autoencoder-based models, Generative Adversarial Networks (GANs), diffusion models, and Large Language Models (LLMs). Finally, we review emerging studies that integrate generative AI with FL-based IDS and discuss open challenges, including synthetic data quality, realistic traffic generation, dual-use adversarial risks, non-IID client distributions, communication-efficient model sharing, federated IDS benchmarking, and domain-specific LLMs for network security.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。