arXiv:2505.04992stat.MLcs.LG2025-05

用预训练大模型生成数据,再通过统计方法筛选优质样本提升预测效果。

Boosting Statistic Learning with Synthetic Data from Pretrained Large Models

  • 基于领域统计方法筛选生成数据,只保留高质量样本
  • 在多种场景下均显著提升预测性能
  • 适合需要数据增强但担心噪声干扰的研究者

生成模型(如Stable Diffusion)的快速发展带来一个关键问题:如何利用其生成的数据提升预测建模?尽管这些模型能生成海量数据,但仅有部分样本能真正改善模型性能。我们提出一种端到端框架,通过领域特定的统计方法生成并系统性过滤合成数据,选择性地整合高质量样本以实现有效数据增强。实验表明,在多种设置下该方法均带来一致的性能提升,凸显了该框架的潜力,同时揭示了生成模型在数据增强中的固有局限性。尽管可生成大量合成数据,但真正能提升性能的比例有限。

原文摘要 · Abstract (English)

The rapid advancement of generative models, such as Stable Diffusion, raises a key question: how can synthetic data from these models enhance predictive modeling? While they can generate vast amounts of datasets, only a subset meaningfully improves performance. We propose a novel end-to-end framework that generates and systematically filters synthetic data through domain-specific statistical methods, selectively integrating high-quality samples for effective augmentation. Our experiments demonstrate consistent improvements in predictive performance across various settings, highlighting the potential of our framework while underscoring the inherent limitations of generative models for data augmentation. Despite the ability to produce large volumes of synthetic data, the proportion that effectively improves model performance is limited.

数据增强生成模型统计方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。