系统分析小样本不平衡问题,揭示现有方法局限性。
A Survey on Small Sample Imbalance Problem: Metrics, Feature Analysis, and Solutions
- 从数据视角构建分析框架,强调特征分布与类间差异
- 实验证明分类器差异影响远大于重采样提升效果
- 适合关注小样本学习与模型可解释性的研究者
小样本不平衡(S&I)问题是机器学习与数据分析中的主要挑战,其特征为样本量少且类别分布极不均衡,导致模型性能下降。此外,类间特征分布模糊进一步加剧分类难度。现有方法多依赖算法启发式策略,缺乏对数据本质特性的深入分析。本文提出系统的分析框架:首先总结不平衡度量与复杂性分析方法,强调可解释基准的重要性;其次综述针对常规、基于复杂性及极端S&I问题的解决方案,揭示不同数据分布下的方法差异;实验表明,尽管重采样被广泛采用,但分类器性能差异显著超过其改进幅度。最后,论文指出开放问题并探讨未来趋势。
原文摘要 · Abstract (English)
The small sample imbalance (S&I) problem is a major challenge in machine learning and data analysis. It is characterized by a small number of samples and an imbalanced class distribution, which leads to poor model performance. In addition, indistinct inter-class feature distributions further complicate classification tasks. Existing methods often rely on algorithmic heuristics without sufficiently analyzing the underlying data characteristics. We argue that a detailed analysis from the data perspective is essential before developing an appropriate solution. Therefore, this paper proposes a systematic analytical framework for the S\&I problem. We first summarize imbalance metrics and complexity analysis methods, highlighting the need for interpretable benchmarks to characterize S&I problems. Second, we review recent solutions for conventional, complexity-based, and extreme S&I problems, revealing methodological differences in handling various data distributions. Our summary finds that resampling remains a widely adopted solution. However, we conduct experiments on binary and multiclass datasets, revealing that classifier performance differences significantly exceed the improvements achieved through resampling. Finally, this paper highlights open questions and discusses future trends.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。