arXiv:2601.16976cs.LG2026-01

用潜在扩散模型生成物联网攻击数据,提升入侵检测准确率。

Latent Diffusion for Internet of Things Attack Data Generation in Intrusion Detection

  • 在潜在空间生成攻击数据,兼顾真实度与多样性
  • 使DDoS和Mirai攻击的F1分数达0.99,超越现有方法
  • 生成速度快25%,适合实际部署的入侵检测系统

基于机器学习的入侵检测系统(ML-based IDS)在保护物联网(IoT)环境方面至关重要,但其性能常因良性流量与攻击流量之间的严重类别不平衡而下降。尽管数据增强被广泛用于缓解此问题,现有方法通常依赖简单的过采样技术或难以同时实现高样本保真度、多样性和计算效率的生成模型。为此,本文提出使用潜在扩散模型(Latent Diffusion Model, LDM)进行物联网入侵检测中的攻击数据增强,并与最先进基线进行全面比较。实验针对分布式拒绝服务(DDoS)、Mirai和中间人(Man-in-the-Middle)三类典型物联网攻击,在下游IDS性能及生成质量上采用分布、依赖关系和多样性指标进行评估。结果表明,使用LDM生成的数据平衡训练集后,显著提升了检测性能,对DDoS和Mirai攻击的F1分数最高达到0.99,且持续优于对比方法。定量与定性分析显示,LDM能有效保持特征依赖关系,生成多样化样本,同时相比直接在数据空间操作的扩散模型,采样时间减少约25%。这些发现表明,潜在扩散是合成物联网攻击数据的有效且可扩展方案,能显著缓解机器学习型入侵检测中因类别不平衡带来的影响。

原文摘要 · Abstract (English)

Intrusion Detection Systems (IDSs) are a key component for protecting Internet of Things (IoT) environments. However, in Machine Learning-based (ML-based) IDSs, performance is often degraded by the strong class imbalance between benign and attack traffic. Although data augmentation has been widely explored to mitigate this issue, existing approaches typically rely on simple oversampling techniques or generative models that struggle to simultaneously achieve high sample fidelity, diversity, and computational efficiency. To address these limitations, we propose the use of a Latent Diffusion Model (LDM) for attack data augmentation in IoT intrusion detection and provide a comprehensive comparison against state-of-the-art baselines. Experiments were conducted on three representative IoT attack types, specifically Distributed Denial-of-Service (DDoS), Mirai, and Man-in-the-Middle, evaluating both downstream IDS performance and intrinsic generative quality using distributional, dependency-based, and diversity metrics. Results show that balancing the training data with LDM-generated samples substantially improves IDS performance, achieving F1-scores of up to 0.99 for DDoS and Mirai attacks and consistently outperforming competing methods. Additionally, quantitative and qualitative analyses demonstrate that LDMs effectively preserve feature dependencies while generating diverse samples and reduce sampling time by approximately 25\% compared to diffusion models operating directly in data space. These findings highlight latent diffusion as an effective and scalable solution for synthetic IoT attack data generation, substantially mitigating the impact of class imbalance in ML-based IDSs for IoT scenarios.

入侵检测扩散模型物联网安全数据生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。