arXiv:2504.07426stat.MEcs.LG2025-04被引 10

用生成模型精准补足数据短板,提升模型泛化能力

Conditional Data Synthesis Augmentation

  • 基于扩散模型生成符合条件分布的合成数据
  • 在稀疏区域提升样本密度,改善数据不平衡问题
  • 适用于表格、文本、图像多模态场景,效果可理论验证

可靠的机器学习与统计分析依赖于多样且分布均匀的训练数据。然而,现实世界数据集往往规模有限,关键子群体代表性不足,导致预测偏差和性能下降,尤其在分类等监督任务中。为应对这一挑战,我们提出条件数据合成增强(CoDSA),一种利用生成模型(如扩散模型)在多模态领域(包括表格、文本和图像数据)中合成高保真数据的新框架。CoDSA生成的合成样本能忠实捕捉原始数据的条件分布,重点覆盖低采样或高关注区域。通过迁移学习,对预训练生成模型进行微调,以提升合成数据的真实感并增加稀疏区域的样本密度。该过程保持了模态间关系,缓解数据不平衡,提升域适应性并增强泛化能力。我们还引入一个理论框架,量化了合成样本数量与目标区域分配对统计精度提升的影响,提供有效性形式保证。大量实验表明,CoDSA在监督与无监督设置下均显著优于非自适应增强策略及现有最优基线。

原文摘要 · Abstract (English)

Reliable machine learning and statistical analysis rely on diverse, well-distributed training data. However, real-world datasets are often limited in size and exhibit underrepresentation across key subpopulations, leading to biased predictions and reduced performance, particularly in supervised tasks such as classification. To address these challenges, we propose Conditional Data Synthesis Augmentation (CoDSA), a novel framework that leverages generative models, such as diffusion models, to synthesize high-fidelity data for improving model performance across multimodal domains including tabular, textual, and image data. CoDSA generates synthetic samples that faithfully capture the conditional distributions of the original data, with a focus on under-sampled or high-interest regions. Through transfer learning, CoDSA fine-tunes pre-trained generative models to enhance the realism of synthetic data and increase sample density in sparse areas. This process preserves inter-modal relationships, mitigates data imbalance, improves domain adaptation, and boosts generalization. We also introduce a theoretical framework that quantifies the statistical accuracy improvements enabled by CoDSA as a function of synthetic sample volume and targeted region allocation, providing formal guarantees of its effectiveness. Extensive experiments demonstrate that CoDSA consistently outperforms non-adaptive augmentation strategies and state-of-the-art baselines in both supervised and unsupervised settings.

数据增强生成模型扩散模型多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。