arXiv:2509.00931stat.MLcs.LG2025-09被引 3

用生成模型和不确定性分析提升小额标注下的信用卡欺诈检测效果

Semi-Supervised Bayesian GANs with Log-Signatures for Uncertainty-Aware Credit Card Fraud Detection

  • 结合贝叶斯GAN与对数签名,处理不规则采样交易序列
  • 在少样本条件下显著提升分类准确率与不确定性量化能力
  • 适合金融风控场景中数据稀缺且需可信预测的场景

我们提出一种新型深度生成半监督框架,用于信用卡欺诈检测,将其建模为时间序列分类任务。随着金融交易数据规模与复杂性增加,传统方法常依赖大量标注数据,难以应对不规则采样频率和长度各异的时间序列。为此,我们扩展条件生成对抗网络以实现针对性数据增强,引入贝叶斯推断获取预测分布并量化不确定性,同时采用对数签名进行交易历史的鲁棒特征编码。设计了一种基于Wasserstein距离的损失函数,使生成样本与真实未标注样本对齐,同时最大化标注数据上的分类精度。在BankSim数据集上评估,不同标注比例下均优于基准模型,在全局统计指标与领域特定指标上表现一致提升。结果表明,基于GAN的半监督学习结合对数签名,对不规则采样时间序列有效,且不确定性感知预测至关重要。

原文摘要 · Abstract (English)

We present a novel deep generative semi-supervised framework for credit card fraud detection, formulated as time series classification task. As financial transaction data streams grow in scale and complexity, traditional methods often require large labeled datasets, struggle with time series of irregular sampling frequencies and varying sequence lengths. To address these challenges, we extend conditional Generative Adversarial Networks (GANs) for targeted data augmentation, integrate Bayesian inference to obtain predictive distributions and quantify uncertainty, and leverage log-signatures for robust feature encoding of transaction histories. We introduce a novel Wasserstein distance-based loss to align generated and real unlabeled samples while simultaneously maximizing classification accuracy on labeled data. Our approach is evaluated on the BankSim dataset, a widely used simulator for credit card transaction data, under varying proportions of labeled samples, demonstrating consistent improvements over benchmarks in both global statistical and domain-specific metrics. These findings highlight the effectiveness of GAN-driven semi-supervised learning with log-signatures for irregularly sampled time series and emphasize the importance of uncertainty-aware predictions.

欺诈检测生成模型不确定性时间序列

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。