提出多层级分布熵,无需原始数据即可解释网络入侵检测。
Multi-Level Distributional Entropy for Explainable Network Intrusion Detection
- 从流量统计中提取三种可解释的熵特征,不依赖原始包数据。
- 仅用熵特征即达0.708-0.989加权F1,性能不降反而暴露隐藏缺陷。
- 适用于需要可解释性与鲁棒性的实际部署场景,尤其适合安全审计。
机器学习网络入侵检测系统依赖聚合流统计量,却丢失分布结构;传统熵度量需原始报文序列,无法用于预聚合流数据。本文提出多层级分布熵(MDE),一种从流级摘要统计中直接提取可解释熵特征的分析框架,涵盖三类:流内高斯微分熵、跨方向Jensen-Shannon散度(JSD)及TCP标志模式香农熵,无需原始包访问或训练数据。在四个基准数据集(NSL-KDD、CICIDS-2017、CICIDS-2018、UNSW-NB15)上采用无泄露折叠局部管道测试,仅使用熵特征即获得0.708–0.989的加权F1,性能媲美常规特征且不降级。完整指标报告揭示了聚合F1掩盖的失败模式:在CICIDS-2018上,F1=0.74时检测率(DR)仅为0.48;对未见攻击家族,F1超过0.998但检测率降至零。在时间漂移下,703K流伪实时重放显示评分排序保持稳定(AUC=0.87),但固定阈值崩溃(DR=0.082),再校准也无法恢复。SHAP折叠稳定性分析(Spearman rho=0.80–0.95)证实熵归因在异构环境中具有可重复性与领域一致性。
原文摘要 · Abstract (English)
Machine learning network intrusion detection systems (IDS) rely on aggregate flow statistics that discard distributional structure, while established entropy measures require raw packet sequences unavailable in pre-aggregated flow datasets. We propose Multi-Level Distributional Entropy (MDE), an analytical framework that derives interpretable entropy features directly from flow-level summary statistics at three levels: within-flow Gaussian differential entropy, cross-directional Jensen-Shannon divergence (JSD), and Transmission Control Protocol (TCP) flag-pattern Shannon entropy, without raw packet access or training data. Across four benchmarks (NSL-KDD, CICIDS-2017, CICIDS-2018, UNSW-NB15) under a leakage-free fold-local pipeline, entropy-only features achieve weighted F1 of 0.708-0.989, matching conventional features without degrading performance. Full operational metric reporting then exposes failure modes that aggregate F1 conceals. On CICIDS-2018, F1=0.74 hides a detection rate (DR) of 0.48, and on held-out attack families F1 exceeds 0.998 while DR falls to zero. Under temporal shift, a pseudo-live replay of 703K flows reveals a threshold-ranking divergence in which score ranking is preserved (AUC=0.87) but fixed thresholds collapse (DR=0.082) and recalibration offers no recovery. SHapley Additive exPlanations (SHAP) fold-stability analysis (Spearman rho=0.80-0.95) confirms that entropy attributions are reproducible and domain-coherent across heterogeneous environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。