用数学积分方法融合网络特征,提升异常检测精度并大幅压缩数据量。
Improving Network Anomaly Detection via Choquet-Integral-Based Feature Aggregation

- 基于柯赫特积分的特征聚合,自适应加权并逐步筛选关键特征。
- 准确率最高提升7%,数据量减少77.5%(214MB→48MB),且不牺牲查全率和查准率。
- 适合带宽受限、特征稀缺的实时入侵检测场景,效果显著优于传统方法。
本文提出一种基于广义柯赫特积分的特征聚合框架,用于提升高维网络流量数据中的异常检测性能。该方法结合自适应加权与增量特征选择,有效缓解特征冗余问题。采用随机森林与XGBoost分类器,在不同特征子集规模下对比原始特征与柯赫特聚合特征的模型表现。结果表明,该聚合方法在不降低精确率和召回率的前提下,准确率最高提升7%,数据量从214MB降至48MB(减少77.5%)。多轮分层重复实验显示,当特征可用性受限时,柯赫特聚合带来的增益具有统计显著性(p < 0.05),凸显其在带宽和特征资源受限的实时入侵检测场景中的适用性。
原文摘要 · Abstract (English)
This work investigates a generalized Choquet-integral-based feature aggregation framework to improve anomaly detection in high-dimensional network traffic data. The approach combines adaptive weighting with incremental feature selection to address feature redundancy. Using Random Forest and XGBoost classifiers, we evaluate models trained with both raw and Choquet-aggregated features under varying feature subset sizes. The proposed aggregation achieves up to $7\%$ higher accuracy while reducing data volume by $77.5\%$ (from $214$~MB to $48$~MB), without degrading precision and recall. Results averaged over multiple stratified repetitions indicate that Choquet-based aggregation yields statistically significant gains ($p < 0.05$) in scenarios with limited feature availability, highlighting its suitability for real-time intrusion detection under bandwidth and feature-availability constraints.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。