用解耦变分自编码器改进不平衡回归的数据生成方法
Disentangled Deep Smoothed Bootstrap for Fair Imbalanced Regression
- 在潜在空间中结合解耦VAE与平滑自助法生成新数据
- 在基准数据集上显著提升不平衡回归的预测性能
- 适合处理表格型数据的不公平分布回归问题
不平衡分布学习是预测建模中的常见且重要挑战,常导致标准算法性能下降。尽管已有多种方法应对该问题,但多数针对分类任务,对回归的关注较少。本文提出一种新方法,用于改善表格数据在不平衡回归(IR)框架下的学习效果。我们利用变分自编码器(VAEs)建模并定义数据分布的潜在表示。然而,VAE在不均衡数据上的效率与其他标准方法类似受限。为此,我们开发了一种创新的数据生成方法,将解耦VAE与在潜在空间中应用的平滑自助法相结合。通过与竞争方法在不平衡回归基准数据集上的数值对比,评估了该方法的有效性。
原文摘要 · Abstract (English)
Imbalanced distribution learning is a common and significant challenge in predictive modeling, often reducing the performance of standard algorithms. Although various approaches address this issue, most are tailored to classification problems, with a limited focus on regression. This paper introduces a novel method to improve learning on tabular data within the Imbalanced Regression (IR) framework, which is a critical problem. We propose using Variational Autoencoders (VAEs) to model and define a latent representation of data distributions. However, VAEs can be inefficient with imbalanced data like other standard approaches. To address this, we develop an innovative data generation method that combines a disentangled VAE with a Smoothed Bootstrap applied in the latent space. We evaluate the efficiency of this method through numerical comparisons with competitors on benchmark datasets for IR.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。