arXiv:2507.16419cs.LG2025-07被引 2

用开源合成数据提升极端不平衡数据的预测准确率

Improving Predictions on Highly Unbalanced Data Using Open Source Synthetic Data Upsampling

  • 用AI生成合成数据填补少数类特征空缺
  • 合成数据使少数类预测准确率显著提升
  • 适合处理医疗、欺诈等稀有事件数据

高度不平衡的表格数据集在欺诈检测、医疗诊断和罕见事件预测等众多领域带来严峻挑战。现实中少数类样本极度稀缺,传统机器学习算法易偏向多数类,导致模型偏差。合成数据可通过生成多样且逼真的新样本缓解少数类欠采样问题。本文对开源合成数据工具MOSTLY AI的Synthetic Data SDK进行基准测试,评估其在混合类型数据上的合成上采样效果。实验对比了使用合成数据、简单重采样及SMOTE-NC方法训练的模型表现。结果表明,合成数据能有效生成填补特征空间稀疏区域的多样化样本,显著提升少数类预测性能。尤其在少数类样本极少的混合类型数据集上,合成数据上采样可稳定获得最优模型表现。

原文摘要 · Abstract (English)

Unbalanced tabular data sets present significant challenges for predictive modeling and data analysis across a wide range of applications. In many real-world scenarios, such as fraud detection, medical diagnosis, and rare event prediction, minority classes are vastly underrepresented, making it difficult for traditional machine learning algorithms to achieve high accuracy. These algorithms tend to favor the majority class, leading to biased models that struggle to accurately represent minority classes. Synthetic data holds promise for addressing the under-representation of minority classes by providing new, diverse, and highly realistic samples. This paper presents a benchmark study on the use of AI-generated synthetic data for upsampling highly unbalanced tabular data sets. We evaluate the effectiveness of an open-source solution, the Synthetic Data SDK by MOSTLY AI, which provides a flexible and user-friendly approach to synthetic upsampling for mixed-type data. We compare predictive models trained on data sets upsampled with synthetic records to those using standard methods, such as naive oversampling and SMOTE-NC. Our results demonstrate that synthetic data can improve predictive accuracy for minority groups by generating diverse data points that fill gaps in sparse regions of the feature space. We show that upsampled synthetic training data consistently results in top-performing predictive models, particularly for mixed-type data sets containing very few minority samples.

合成数据不平衡数据表格数据AI生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。