arXiv:2510.08123stat.MLcs.LG2025-10被引 4

发现合成数据的协方差匹配比均值匹配更能提升分类器性能。

High-dimensional Analysis of Synthetic Data Selection

  • 从高维回归视角分析,协方差偏移影响泛化误差,均值偏移则不影响。
  • 在特定条件下,匹配目标分布的协方差可达到最优性能。
  • 该结论在深度神经网络和生成模型中依然有效,适合数据增强场景。

尽管生成模型取得进展,其生成的合成数据能否真正提升分类器性能仍存疑。现有方法多依赖“合成数据应接近真实分布”等经验原则,但具体哪些属性影响泛化误差尚不明确。本文从高维回归角度研究该问题:理论证明,对线性模型而言,目标分布与合成数据之间的协方差偏移会影响泛化误差,而均值偏移则无影响;在某些设定下,匹配目标分布的协方差即为最优。令人惊讶的是,这些线性模型的理论洞察可推广至深度神经网络与生成模型。实验表明,协方差匹配策略在多种训练范式、架构、数据集及生成模型下,均优于多项近期合成数据选择方法。

原文摘要 · Abstract (English)

Despite the progress in the development of generative models, their usefulness in creating synthetic data that improve prediction performance of classifiers has been put into question. Besides heuristic principles such as "synthetic data should be close to the real data distribution", it is actually not clear which specific properties affect the generalization error. Our paper addresses this question through the lens of high-dimensional regression. Theoretically, we show that, for linear models, the covariance shift between the target distribution and the distribution of the synthetic data affects the generalization error but, surprisingly, the mean shift does not. Furthermore we prove that, in some settings, matching the covariance of the target distribution is optimal. Remarkably, the theoretical insights from linear models carry over to deep neural networks and generative models. We empirically demonstrate that the covariance matching procedure (matching the covariance of the synthetic data with that of the data coming from the target distribution) performs well against several recent approaches for synthetic data selection, across training paradigms, architectures, datasets and generative models used for augmentation.

合成数据协方差匹配泛化误差生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。