arXiv:2601.01927cs.LGcs.AI2026-01

证明了SMOTE生成样本会收敛到原始数据分布,为数据增强提供理论支撑。

Theoretical Convergence of SMOTE-Generated Samples

  • 从概率和均值角度严格证明合成样本收敛于原始数据。
  • 当原始数据有界时,合成样本均值收敛速度更快。
  • 邻居数越小,收敛越快,可指导实际应用参数设置。

不平衡数据广泛影响医疗、网络安全等机器学习应用。作为主流解决方法之一,SMOTE的理论基础亟需验证。本文对SMOTE的收敛性进行严格理论分析:证明合成随机变量Z在概率意义下收敛于原始随机变量X;当X有界时,进一步证明其在均值意义下收敛;并表明邻居数越小,收敛速度越快,为实践提供可操作建议。数值实验使用真实与合成数据验证了理论结果。本工作为数据增强技术提供了基础理论支持,适用于更广泛的场景。

原文摘要 · Abstract (English)

Imbalanced data affects a wide range of machine learning applications, from healthcare to network security. As SMOTE is one of the most popular approaches to addressing this issue, it is imperative to validate it not only empirically but also theoretically. In this paper, we provide a rigorous theoretical analysis of SMOTE's convergence properties. Concretely, we prove that the synthetic random variable Z converges in probability to the underlying random variable X. We further prove a stronger convergence in mean when X is compact. Finally, we show that lower values of the nearest neighbor rank lead to faster convergence offering actionable guidance to practitioners. The theoretical results are supported by numerical experiments using both real-life and synthetic data. Our work provides a foundational understanding that enhances data augmentation techniques beyond imbalanced data scenarios.

SMOTE数据增强理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。