用优化方法自动选最优分类器组合,提升罕见事件检测准确率。
Rare Event Detection in Imbalanced Multi-Class Datasets Using an Optimal MIP-Based Ensemble Weighting Approach
- 基于混合整数规划,为每个类别动态分配分类器权重。
- 相比现有方法,平衡准确率平均提升4.53%,其他指标均增4.6%以上。
- 适合需要高精度罕见事件识别的工业安全系统场景。
针对关键网络物理系统中罕见事件检测面临的多类不平衡数据挑战,本文提出一种最优、高效且可适应的混合整数规划(MIP)集成加权方法。该方法在细粒度类别基础上利用分类器集成的多样化能力,通过弹性网络正则化优化分类器-类别对的权重,增强模型鲁棒性与泛化能力,并能无缝、最优地从给定集合中选出预设数量的分类器。我们在多个代表性数据集上,使用合适评估指标,对比了六种主流加权方案,在不同集成规模下进行实验。结果表明,所提MIP方法优于所有现有方法:在各类数据集与集成规模下,平衡准确率提升0.99%至7.31%,平均提升4.53%;宏平均精确率、召回率和F1值分别平均提高4.63%、4.60%和4.61%,同时保持计算效率。
原文摘要 · Abstract (English)
To address the challenges of imbalanced multi-class datasets typically used for rare event detection in critical cyber-physical systems, we propose an optimal, efficient, and adaptable mixed integer programming (MIP) ensemble weighting scheme. Our approach leverages the diverse capabilities of the classifier ensemble on a granular per class basis, while optimizing the weights of classifier-class pairs using elastic net regularization for improved robustness and generalization. Additionally, it seamlessly and optimally selects a predefined number of classifiers from a given set. We evaluate and compare our MIP-based method against six well-established weighting schemes, using representative datasets and suitable metrics, under various ensemble sizes. The experimental results reveal that MIP outperforms all existing approaches, achieving an improvement in balanced accuracy ranging from 0.99% to 7.31%, with an overall average of 4.53% across all datasets and ensemble sizes. Furthermore, it attains an overall average increase of 4.63%, 4.60%, and 4.61% in macro-averaged precision, recall, and F1-score, respectively, while maintaining computational efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。