arXiv:2412.14527stat.MLcs.LG2024-12被引 5

用信息论选关键样本,解决数据不均衡问题。

Statistical Undersampling with Mutual Information and Support Points

  • 基于互信息与支持点优化选择代表性样本
  • 在多个任务中平衡准确率优于传统方法
  • 适合处理数据分布不均的分类场景

大规模数据集中的类别不平衡与分布差异给分类任务带来显著挑战,常导致模型偏差及少数类预测性能差。本文提出两种新型欠采样方法:基于互信息的分层简单随机抽样与支持点优化。这些方法优先选取具有代表性的数据,有效减少信息损失。在多个分类任务上的实验结果表明,所提方法优于传统欠采样技术,实现了更高的平衡分类准确率。研究揭示了将统计学概念与机器学习结合以应对实际应用中类别不平衡问题的潜力。

原文摘要 · Abstract (English)

Class imbalance and distributional differences in large datasets present significant challenges for classification tasks machine learning, often leading to biased models and poor predictive performance for minority classes. This work introduces two novel undersampling approaches: mutual information-based stratified simple random sampling and support points optimization. These methods prioritize representative data selection, effectively minimizing information loss. Empirical results across multiple classification tasks demonstrate that our methods outperform traditional undersampling techniques, achieving higher balanced classification accuracy. These findings highlight the potential of combining statistical concepts with machine learning to address class imbalance in practical applications.

欠采样类别不平衡信息论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。