纠正了过采样方法选择的常见误区,发现数据可分性比不平衡率更重要。
Beyond Imbalance Ratio: Data Characteristics as Critical Moderators of Oversampling Method Selection
- 通过生成高斯混合数据控制变量,系统测试不同不平衡率下的过采样效果
- 发现不平衡率与过采样收益呈弱到中度负相关,数据可分性影响更大
- 提出'情境决定选择'框架,指导实际应用中科学选法
主流的不平衡率阈值范式认为不平衡率(IR)越高,过采样效果越好,但这一假设缺乏受控实验支持。我们进行了12次受控实验(共超过100个数据变体),通过算法生成高斯混合数据,在保持数据特征(类别可分性、聚类结构)恒定的前提下系统性地改变IR。另两个验证实验考察了上限效应和指标依赖性。所有方法在来自OpenML的17个真实数据集上进行评估。在控制混杂变量后,IR与过采样收益呈现弱至中度负相关。类别可分性成为更强的调节因子,其对方法有效性解释方差显著高于仅考虑IR。我们提出‘情境决定选择’框架,整合IR、类别可分性和聚类结构,为实践者提供基于证据的方法选择标准。
原文摘要 · Abstract (English)
The prevailing IR-threshold paradigm posits a positive correlation between imbalance ratio (IR) and oversampling effectiveness, yet this assumption remains empirically unsubstantiated through controlled experimentation. We conducted 12 controlled experiments (N > 100 dataset variants) that systematically manipulated IR while holding data characteristics (class separability, cluster structure) constant via algorithmic generation of Gaussian mixture datasets. Two additional validation experiments examined ceiling effects and metric-dependence. All methods were evaluated on 17 real-world datasets from OpenML. Upon controlling for confounding variables, IR exhibited a weak to moderate negative correlation with oversampling benefits. Class separability emerged as a substantially stronger moderator, accounting for significantly more variance in method effectiveness than IR alone. We propose a 'Context Matters' framework that integrates IR, class separability, and cluster structure to provide evidence-based selection criteria for practitioners.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。