arXiv:2602.06955cs.LG2026-02

优化可解释模型提升信用卡欺诈检测准确率

Improving Credit Card Fraud Detection with an Optimized Explainable Boosting Machine

  • 用正交实验法优化数据预处理与模型参数顺序
  • 在基准数据集上达到0.983的ROC-AUC,优于多个主流模型
  • 兼顾高精度与可解释性,适合金融风控场景

信用卡欺诈检测面临类别不平衡的核心挑战,直接影响实际金融系统的预测可靠性。本文提出一种基于可解释提升机(EBM)的优化流程,通过系统化的超参数调优、特征选择与预处理改进,克服传统采样方法引入偏差或信息损失的问题。采用田口方法优化数据缩放器序列与模型超参数,实现稳定、可复现的性能提升。在基准信用卡数据集上的实验显示,该方法获得0.983的ROC-AUC,优于先前EBM基线(0.975),并超越逻辑回归、随机森林、XGBoost和决策树模型。结果表明,可解释机器学习与数据驱动优化结合,有助于推动可信金融欺诈分析的发展。

原文摘要 · Abstract (English)

Addressing class imbalance is a central challenge in credit card fraud detection, as it directly impacts predictive reliability in real-world financial systems. To overcome this, the study proposes an enhanced workflow based on the Explainable Boosting Machine (EBM)-a transparent, state-of-the-art implementation of the GA2M algorithm-optimized through systematic hyperparameter tuning, feature selection, and preprocessing refinement. Rather than relying on conventional sampling techniques that may introduce bias or cause information loss, the optimized EBM achieves an effective balance between accuracy and interpretability, enabling precise detection of fraudulent transactions while providing actionable insights into feature importance and interaction effects. Furthermore, the Taguchi method is employed to optimize both the sequence of data scalers and model hyperparameters, ensuring robust, reproducible, and systematically validated performance improvements. Experimental evaluation on benchmark credit card data yields an ROC-AUC of 0.983, surpassing prior EBM baselines (0.975) and outperforming Logistic Regression, Random Forest, XGBoost, and Decision Tree models. These results highlight the potential of interpretable machine learning and data-driven optimization for advancing trustworthy fraud analytics in financial systems.

欺诈检测可解释模型优化算法金融风控

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。