用熵值特征提升网络异常检测准确率,效果好且开销小。
On the Impact of Entropy-based Features

- 在传统统计特征外加入熵值,捕捉流量变异性的新方法。
- 实验显示分类性能稳定提升,尤其在高变异性场景下误判减少。
- 适合对轻量级和可解释性要求高的实际部署场景。
随着流量模式日益多样和复杂,传统统计特征难以充分刻画网络行为。本文探索将熵作为补充特征用于监督式网络流量分类,通过熵量化选定流量属性的变异性,与常规特征互补而非替代。将熵特征集成至标准机器学习流程,在公开入侵检测数据集上进行对比实验,结果显示加入熵特征后分类性能持续提升,计算开销极低。混淆矩阵分析表明,高变异性流量场景下的误分类显著减少。结果表明,熵特征是一种简单实用的方法,可有效增强现有异常检测系统。该方法在注重轻量特征工程与可解释性的场景中尤为适用。
原文摘要 · Abstract (English)
Network anomaly detection is increasingly challenging due to the growing diversity and variability of traffic patterns, which are not always well captured by traditional statistical features. In this work, we explore the use of entropy as an additional feature to support supervised network traffic classification. The main idea is to use entropy to represent variability in selected traffic attributes, complementing conventional descriptors rather than replacing them. We integrate the entropy-based feature into a standard machine learning pipeline and evaluate its impact through a direct comparison between models trained with and without this feature. Experiments conducted on a public intrusion detection dataset show consistent improvements in classification performance, while the additional computational cost remains low. The analysis of confusion matrices indicates a reduction in misclassifications, especially in traffic scenarios with higher variability. Overall, the results suggest that entropy-based features offer a simple and practical way to enhance existing anomaly detection pipelines. This approach is particularly attractive in settings where lightweight feature engineering and interpretability are important, making entropy a useful complement to commonly used traffic features.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。