arXiv:2510.20472stat.MLcs.LG2025-10NeurIPS被引 2

为SMOTE等数据增强方法提供理论支撑,揭示其在不平衡分类中的风险控制机制。

Concentration and excess risk bounds for imbalanced classification with synthetic oversampling

  • 基于核方法构建合成样本的统一集中性边界
  • 给出非参数分类器在合成数据上的泛化误差上界
  • 提出参数调优指导,适合做机器学习理论研究者

使用SMOTE及其变体对少数类样本进行合成扩充,是解决不平衡分类问题的主流策略。尽管该方法在实践中表现良好,但其理论基础仍不充分。本文建立了一个理论框架,分析基于合成数据训练的分类器行为。首先推导出合成少数类样本经验风险与真实少数类分布总体风险之间差异的统一集中性界;随后为基于核方法的分类器在合成数据上训练时提供非参数化的过拟合风险保证。这些结果为SMOTE及下游学习算法的参数调优提供了实用指导。数值实验验证并支持了理论发现。

原文摘要 · Abstract (English)

Synthetic oversampling of minority examples using SMOTE and its variants is a leading strategy for addressing imbalanced classification problems. Despite the success of this approach in practice, its theoretical foundations remain underexplored. We develop a theoretical framework to analyze the behavior of SMOTE and related methods when classifiers are trained on synthetic data. We first derive a uniform concentration bound on the discrepancy between the empirical risk over synthetic minority samples and the population risk on the true minority distribution. We then provide a nonparametric excess risk guarantee for kernel-based classifiers trained using such synthetic data. These results lead to practical guidelines for better parameter tuning of both SMOTE and the downstream learning algorithm. Numerical experiments are provided to illustrate and support the theoretical findings

分类SMOTE理论分析不平衡数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。