arXiv:2509.23240cs.LG2025-09

用扩散模型生成少数类特征,提升高维数据回归的准确性

More Data or Better Algorithms: Latent Diffusion Augmentation for Deep Imbalanced Regression

  • 在隐空间用条件扩散模型优先生成少数类样本
  • 在三个基准上显著改善少数区域预测效果
  • 适用于图像、文本等高维数据,效率高

许多现实世界的回归任务中数据分布严重偏斜,模型主要学习多数样本而难以准确预测少数标签。尽管不平衡分类研究充分,但不平衡回归仍相对未被探索。深度不平衡回归(DIR)指输入为高维非结构化数据的情形。虽已有针对表格数据的不平衡回归数据级方法,但当前深度不平衡回归缺乏适用于高维数据的数据级解决方案,主要依赖算法修改。为此,我们提出LatentDiff框架,利用带优先级生成的条件扩散模型,在隐表示空间合成高质量特征。LatentDiff计算高效,适用于多种数据模态,包括图像、文本等高维输入。在三个DIR基准上的实验表明,该方法在保持整体精度的同时,显著提升了少数区域的预测性能。

原文摘要 · Abstract (English)

In many real-world regression tasks, the data distribution is heavily skewed, and models learn predominantly from abundant majority samples while failing to predict minority labels accurately. While imbalanced classification has been extensively studied, imbalanced regression remains relatively unexplored. Deep imbalanced regression (DIR) represents cases where the input data are high-dimensional and unstructured. Although several data-level approaches for tabular imbalanced regression exist, deep imbalanced regression currently lacks dedicated data-level solutions suitable for high-dimensional data and relies primarily on algorithmic modifications. To fill this gap, we propose LatentDiff, a novel framework that uses conditional diffusion models with priority-based generation to synthesize high-quality features in the latent representation space. LatentDiff is computationally efficient and applicable across diverse data modalities, including images, text, and other high-dimensional inputs. Experiments on three DIR benchmarks demonstrate substantial improvements in minority regions while maintaining overall accuracy.

不平衡回归扩散模型数据增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。