arXiv:2409.09792cs.LG2024-09被引 2

用自学习提升金融风险数据质量,改善少数类预测效果。

Enhancing Data Quality through Self-learning on Imbalanced Financial Risk Data

  • 三步法:生成少数类样本、二值反馈过滤、伪标签自学习
  • 在6个数据集上显著提升少数类校准性能
  • 适合需要高精度风险预警的金融风控场景

在信用违约预测和欺诈检测等金融风险领域,准确识别高风险样本至关重要,因其影响巨大。尽管机器学习模型广泛应用,但其性能常受限于高质量数据的稀缺与多样性不足。问题源于数据集中的小样本量、高标注成本及严重类别不平衡,阻碍模型有效学习并准确预测关键事件。本研究提出TriEnhance方法,通过三步优化现有金融风险数据集:(1)针对性生成少数类合成样本;(2)利用二值反馈进行样本筛选以精炼数据;(3)基于伪标签的自学习机制。在六个基准数据集上的实验表明,该方法显著提升了少数类的校准性能,是构建更稳健金融风险预测系统的关键。

原文摘要 · Abstract (English)

In the financial risk domain, particularly in credit default prediction and fraud detection, accurate identification of high-risk class instances is paramount, as their occurrence can have significant economic implications. Although machine learning models have gained widespread adoption for risk prediction, their performance is often hindered by the scarcity and diversity of high-quality data. This limitation stems from factors in datasets such as small risk sample sizes, high labeling costs, and severe class imbalance, which impede the models' ability to learn effectively and accurately forecast critical events. This study investigates data pre-processing techniques to enhance existing financial risk datasets by introducing TriEnhance, a straightforward technique that entails: (1) generating synthetic samples specifically tailored to the minority class, (2) filtering using binary feedback to refine samples, and (3) self-learning with pseudo-labels. Our experiments across six benchmark datasets reveal the efficacy of TriEnhance, with a notable focus on improving minority class calibration, a key factor for developing more robust financial risk prediction systems.

金融风控数据增强自学习少数类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。