arXiv:2608.10349cs.CRcs.LG2026-08

评估物联网入侵检测的解释成本、稳定性和实用性,超越单纯准确率。

Beyond Detection Accuracy: Measuring Explanation Cost, Stability, and Utility for Resource-Aware IoT Intrusion Detection

  • 联合评估预测效果、解释开销、局部稳定性与选择性解释。
  • 随机森林解释更稳定,但XGBoost在计算效率上快约476倍。
  • 按需生成解释可节省28%-32%算力,适合资源受限场景。

机器学习入侵检测研究通常强调预测准确率,而将解释生成视为无成本的后处理步骤。本研究首次联合评估二分类物联网(IoT)入侵检测中的预测有效性、解释成本、局部解释稳定性及选择性解释。基于精确39特征哈希、非有限值处理、特征去重、保守标签冲突剔除和确定性哈希级划分,构建了安全泄漏的CICIoT2023数据集。对逻辑回归、决策树、随机森林和XGBoost在自然分布与平衡测试集上进行了评估。测量了TreeSHAP的计算成本,通过预测保持扰动评估稳定性,并使用校准验证策略分配解释负载。XGBoost展现出最强的整体预测性能,随机森林则具有最低误报率。在5,000样本下,随机森林的TreeSHAP耗时700.759秒,而XGBoost仅需1.471秒。随机森林在基础解释稳定性上表现最优;XGBoost虽保持高排名与方向一致性,但顶级特征更换频繁,归因幅度漂移明显。在平衡测试集中,约90%的误报解释覆盖率可实现28-32%的计算节省,约95%覆盖率可节省15-23%。但在攻击密集的自然分布下,节省幅度显著降低。结果表明,实用的可解释物联网入侵检测需综合考虑预测质量、解释成本、局部稳定性、负载分布及选择性调用,而非仅依赖检测准确率。

原文摘要 · Abstract (English)

Machine-learning intrusion-detection studies commonly emphasize predictive accuracy while treating explanation generation as a computationally free post-processing step. This study jointly evaluates predictive effectiveness, explanation cost, local explanation stability, and selective explanation for binary Internet of Things (IoT) intrusion detection. A leakage-safe CICIoT2023 corpus was constructed using exact 39-feature hashes, non-finite-value handling, exact-feature deduplication, conservative label-collision removal, and deterministic hash-level partitioning. Logistic Regression, Decision Tree, Random Forest, and XGBoost were evaluated on natural and balanced test distributions. TreeSHAP cost was measured, stability was assessed under prediction-preserving perturbations, and validation-calibrated policies were used to allocate explanation workload. XGBoost provided the strongest overall predictive profile, while Random Forest produced the lowest false-positive rate. At 5,000 samples, TreeSHAP required 700.759 s for Random Forest and 1.471 s for XGBoost. Random Forest showed the strongest overall base-level explanation stability; XGBoost retained high rank and directional consistency but showed greater top-feature turnover and attribution-magnitude drift. On the balanced test, about 90% false-negative explanation coverage permitted 28-32% compute savings, while about 95% coverage permitted 15-23% savings. Savings were much smaller under the attack-heavy natural prevalence. These results show that operationally useful explainable IoT intrusion detection depends on predictive quality, explanation cost, local stability, workload prevalence, and selective invocation rather than detection accuracy alone.

入侵检测解释性AI资源约束物联网安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。