arXiv:2509.05766cs.LGstat.ML2025-09

用自编码器增强异常检测模型的准确性和可解释性。

Ensemble of Precision-Recall Curve (PRC) Classification Trees with Autoencoders

  • 结合自编码器与精确率-召回率树,学习低维特征并分类异常
  • 在多个基准数据集上表现优于现有方法,提升准确率与扩展性
  • 适合对可解释性要求高的安全、金融等高风险场景

异常检测在网络安全、入侵检测和反欺诈等关键应用中至关重要,需快速识别异常模式。该领域常面临极端类别不平衡和高维数据两大挑战。此前我们提出基于精确率-召回率曲线(PRC)的分类树及其集成方法PRC随机森林(PRC-RF)。本文在此基础上,提出一种融合自编码器与PRC-RF的混合框架,利用自编码器学习紧凑的潜在表示,同时应对两类挑战。在多个基准数据集上的大量实验表明,所提Autoencoder-PRC-RF模型在准确性、可扩展性和可解释性方面均优于已有方法,展现出在高风险异常检测任务中的巨大潜力。

原文摘要 · Abstract (English)

Anomaly detection underpins critical applications from network security and intrusion detection to fraud prevention, where recognizing aberrant patterns rapidly is indispensable. Progress in this area is routinely impeded by two obstacles: extreme class imbalance and the curse of dimensionality. To combat the former, we previously introduced Precision-Recall Curve (PRC) classification trees and their ensemble extension, the PRC Random Forest (PRC-RF). Building on that foundation, we now propose a hybrid framework that integrates PRC-RF with autoencoders, unsupervised machine learning methods that learn compact latent representations, to confront both challenges simultaneously. Extensive experiments across diverse benchmark datasets demonstrate that the resulting Autoencoder-PRC-RF model achieves superior accuracy, scalability, and interpretability relative to prior methods, affirming its potential for high-stakes anomaly-detection tasks.

异常检测自编码器可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。