arXiv:2606.26053stat.MLcs.LG2026-06

合成数据增强在特定条件下能提升不平衡分类性能,否则可能无效甚至有害。

When Does Synthetic Data Augmentation Improve Score-Based Imbalanced Classification?

论文配图:When Does Synthetic Data Augmentation Improve Score-Based Imbalanced Classification?
图 1 · 摘自论文原文
  • 区分增广的权重效应与分布偏差,揭示其理论边界
  • 在模型正确设定时,原始方法已最优,增广仅能降方差
  • 模型误设时,增广可纠正排序错误,带来非单调提升

合成数据增强广泛用于缓解类别不平衡问题,但其对基于得分的分类指标的理论影响仍不明确。本文构建框架,分析合成少数类样本在阈值集成与优化指标(如AUROC、AUPRC、最佳阈值平衡准确率、最佳阈值F1)上的改进条件。将增广效果分解为有效类别权重变化与合成分布和真实分布间的差异。在模型正确设定下,原始估计器已逼近似然比排序,该排序对所考虑指标为总体最优,因此增广无法带来根本性总体改进,仅可能减少有限样本方差,且可能因合成分布误差引入偏差。我们进一步建立极小极大下界,表明原始估计器在正确设定情形下已达到最优指标遗憾率。而在模型误设时,增广可通过改变有效类别平衡,修正由原始不平衡目标引起的受限类投影错误,从而产生定性不同作用。我们提供显式改进边界,量化近似误差、有限样本估计误差与合成分布误差的作用。模拟研究验证了理论,显示在正确设定下增益有限,而在误设下存在非单调但显著的提升。

原文摘要 · Abstract (English)

Synthetic data augmentation is widely used to mitigate class imbalance, but its theoretical effects on score-based classification remain poorly understood. This paper develops a framework for characterizing when synthetic minority augmentation can improve threshold-integrated and threshold-optimized metrics, including AUROC, AUPRC, best-threshold balanced accuracy, and best-threshold \(\F_1\) score. We separate the effect of augmentation into two components: a change in effective class weighting and a discrepancy between the synthetic and true minority distributions. Under well-specified score models, the raw estimator already targets the likelihood-ratio ordering, which is population-optimal for the metrics considered. Consequently, augmentation cannot provide a fundamental population-level improvement beyond possible finite-sample variance reduction, and may introduce additional bias through synthetic distributional error. We further establish minimax lower bounds showing that the raw estimator already achieves the optimal metric-regret rate in the well-specified regime. Under misspecification, however, augmentation can play a qualitatively different role: by changing the effective class balance, it can alter the restricted-class projection and correct ranking errors induced by the raw imbalanced objective. We provide explicit improvement bounds quantifying the roles of approximation error, finite-sample estimation error, and synthetic distributional error. Simulation studies corroborate the theory, demonstrating limited gains under well-specification and nontrivial but nonmonotone improvements under misspecification.

不平衡分类合成数据理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。