arXiv:2504.09147cs.LG2025-04被引 2

改进SMOTE生成更精准的少数类样本,提升不平衡数据分类效果。

Kernel-Based Enhanced Oversampling Method for Imbalanced Classification

  • 用凸组合与核加权生成合成样本,更贴近真实分布。
  • 在多个数据集上F1、G-mean、AUC均优于现有方法。
  • 适合医疗诊断等少数类关键的场景使用。

本文提出一种新型过采样技术,用于改善不平衡数据集上的分类性能。该方法在传统SMOTE基础上引入凸组合与核函数加权机制,生成更能代表少数类特征的合成样本。通过在多个真实世界数据集上的实验验证,新方法在F1-score、G-mean和AUC指标上均优于现有技术,为不平衡分类任务提供了一种稳健有效的解决方案。

原文摘要 · Abstract (English)

This paper introduces a novel oversampling technique designed to improve classification performance on imbalanced datasets. The proposed method enhances the traditional SMOTE algorithm by incorporating convex combination and kernel-based weighting to generate synthetic samples that better represent the minority class. Through experiments on multiple real-world datasets, we demonstrate that the new technique outperforms existing methods in terms of F1-score, G-mean, and AUC, providing a robust solution for handling imbalanced datasets in classification tasks.

过采样不平衡分类SMOTE样本生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。