arXiv:2506.01486cs.LG2025-06被引 1

针对回归任务中的数据不平衡问题,提出新方法提升罕见样本预测效果。

Model-agnostic Mitigation Strategies of Data Imbalance for Regression

  • 引入密度-距离与密度比相关性函数,量化数据重要性。
  • crbSMOGN方法在神经网络上显著优于现有技术,尤其对稀有样本。
  • 构建混合模型集成,缓解罕见与常见样本间的性能权衡。

数据不平衡在回归任务中持续存在,导致模型性能偏差并降低预测可靠性,尤其影响对罕见事件的预测。本文综述了基于采样的方法与代价敏感学习的最新进展,并提出新型缓解策略。为更好评估数据重要性,引入密度-距离和密度比相关性函数,融合数据经验频率与领域偏好,增强可解释性。进一步提出cSMOGN与crbSMOGN两种先进缓解技术,改进现有采样方法。在10个合成数据集与42个真实世界数据集上,使用神经网络、XGBoost与随机森林进行量化评估。结果表明,多数策略虽提升稀有样本表现,却损害常见样本性能,且该权衡随稀有样本优化程度加剧。通过构建一个经不平衡缓解训练的模型与一个未处理的模型的集成,可有效缓解此问题。关键发现显示,结合密度比相关性函数的crbSMOGN在神经网络上表现最优,超越现有最佳方法。

原文摘要 · Abstract (English)

Data imbalance persists as a pervasive challenge in regression tasks, introducing bias in model performance and undermining predictive reliability. This is particularly detrimental in applications aimed at predicting rare events that fall outside of the domain of the bulk of the training data. In this study, we review the current state-of-the-art regarding sampling-based methods and cost-sensitive learning. Additionally, we propose novel approaches to mitigate model bias. To better assess the importance of data, we introduce the density-distance and density-ratio relevance functions, which effectively integrate empirical frequency of data with domain-specific preferences, offering enhanced interpretability for end-users. Furthermore, we present advanced mitigation techniques (cSMOGN and crbSMOGN), which build upon and improve existing sampling methods. In a quantitative evaluation, we benchmark state-of-the-art methods on 10 synthetic and 42 real-world datasets, using neural networks, XGBoosting trees and Random Forest models. Our analysis shows that while most strategies improve performance on rare samples, they degrade it on frequent ones. The trade-off becomes larger the more the performance on rare samples is increased. However, to reduce this effect we demonstrate that constructing an ensemble of models -- one trained with imbalance mitigation and another without -- can be used. The key findings underscore the superior performance of our novel crbSMOGN sampling technique with the density-ratio relevance function for neural networks, outperforming state-of-the-art methods.

回归分析数据不平衡采样优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。