针对异常数据生成难题,zGAN可合成真实且带异常的表格数据。
zGAN: An Outlier-focused Generative Adversarial Network For Realistic Synthetic Data Generation
- 基于真实数据协方差生成异常样本,保持特征间相关性。
- 在金融风控等场景中提升模型性能,合成数据逼近真实分布。
- 适合需要增强异常样本的风控、检测类任务使用。
“黑天鹅”事件对传统机器学习模型性能构成根本挑战。后疫情环境下异常情况频发,促使研究者探索合成数据作为真实数据的补充。本文提出zGAN模型,用于生成带有异常特征的合成表格数据。该模型在二分类任务中表现出色,能生成逼真的合成数据并提升模型性能。zGAN的独特之处在于其能复现真实数据中特征间的相关性,并基于真实或人工生成的协方差矩阵生成异常样本。这一方法可用于建模复杂经济事件,增强预测模型训练中的异常样本,支持异常检测、处理或剔除。研究在私有(金融信贷风险)和公开数据集上进行了实验与对比分析。
原文摘要 · Abstract (English)
The phenomenon of "black swans" has posed a fundamental challenge to performance of classical machine learning models. The perceived rise in frequency of outlier conditions, especially in post-pandemic environment, has necessitated exploration of synthetic data as a complement to real data in model training. This article provides a general overview and experimental investigation of the zGAN model architecture developed for the purpose of generating synthetic tabular data with outlier characteristics. The model is put to test in binary classification environments and shows promising results on realistic synthetic data generation, as well as uplift capabilities vis-à-vis model performance. A distinctive feature of zGAN is its enhanced correlation capability between features in the generated data, replicating correlations of features in real training data. Furthermore, crucial is the ability of zGAN to generate outliers based on covariance of real data or synthetically generated covariances. This approach to outlier generation enables modeling of complex economic events and augmentation of outliers for tasks such as training predictive models and detecting, processing or removing outliers. Experiments and comparative analyses as part of this study were conducted on both private (credit risk in financial services) and public datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。