arXiv:2506.02703cs.LGcs.AI2025-06被引 17

简单模型因评估漏洞竟能骗过学术界,暴露出风控研究的严重方法缺陷。

Data Leakage and Deceptive Performance: A Critical Examination of Credit Card Fraud Detection Methodologies

  • 用错误数据预处理制造信息泄露,让模型提前‘看到’未来数据
  • 仅靠优化召回率就达99.9%但精确率极低,结果虚高不可信
  • 适合关注模型评估规范的研究者与实操工程师参考

本研究批判性审视信用卡欺诈检测研究中的方法论严谨性,揭示基本评估缺陷如何掩盖算法复杂性。通过故意采用不当评估流程,我们证明即使简单模型在违反基本方法原则时也能产生看似出色的性能。分析发现当前方法存在四大问题:(1) 预处理序列中普遍存在数据泄露;(2) 方法报告故意模糊不清;(3) 交易数据缺乏充分的时间验证;(4) 通过牺牲精确率过度优化召回率。案例研究显示,一个带有数据泄露的最小神经网络架构,在未遵循正确评估标准的情况下,仍能实现99.9%的召回率,超越文献中许多复杂模型。这些发现表明,在欺诈检测研究中,正确的评估方法比模型复杂度更为关键。研究警示:方法严谨性必须优先于模型设计,对机器学习领域的研究实践具有广泛启示。

原文摘要 · Abstract (English)

This study critically examines the methodological rigor in credit card fraud detection research, revealing how fundamental evaluation flaws can overshadow algorithmic sophistication. Through deliberate experimentation with improper evaluation protocols, we demonstrate that even simple models can achieve deceptively impressive results when basic methodological principles are violated. Our analysis identifies four critical issues plaguing current approaches: (1) pervasive data leakage from improper preprocessing sequences, (2) intentional vagueness in methodological reporting, (3) inadequate temporal validation for transaction data, and (4) metric manipulation through recall optimization at precision's expense. We present a case study showing how a minimal neural network architecture with data leakage outperforms many sophisticated methods reported in literature, achieving 99.9\% recall despite fundamental evaluation flaws. These findings underscore that proper evaluation methodology matters more than model complexity in fraud detection research. The study serves as a cautionary example of how methodological rigor must precede architectural sophistication, with implications for improving research practices across machine learning applications.

欺诈检测评估漏洞数据泄露模型可信

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。