提出合成数据与真实数据的最优比例理论,避免过量合成数据导致性能下降。
Beyond Real Data: Synthetic Data through the Lens of Regularization
- 用算法稳定性推导泛化误差上界,量化合成与真实数据的权衡。
- 发现合成数据比例呈U形曲线,存在最小测试误差的最优比例。
- 适用于医疗影像等数据稀缺场景,指导实际训练中的数据混合策略。
当真实数据稀缺时,合成数据可提升模型泛化能力,但过度依赖可能引入分布不匹配从而降低性能。本文提出一个学习理论框架,量化合成数据与真实数据之间的权衡。该方法利用算法稳定性推导泛化误差上界,将最小化期望测试误差的最优合成-真实数据比例表征为真实与合成分布间Wasserstein距离的函数。我们在混合数据下的核岭回归设置中验证该框架,提供独立感兴趣的详细分析。理论预测合成数据比例存在最优值,导致测试误差随合成数据比例呈现U形变化。我们在CIFAR-10和临床脑MRI数据集上实证验证了该预测。理论进一步扩展至领域自适应场景,表明在有限源数据基础上合理融合合成目标数据,可缓解领域偏移并增强泛化能力。最后,我们为域内与域外场景提供了实用的数据应用指导。
原文摘要 · Abstract (English)
Synthetic data can improve generalization when real data is scarce, but excessive reliance may introduce distributional mismatches that degrade performance. In this paper, we present a learning-theoretic framework to quantify the trade-off between synthetic and real data. Our approach leverages algorithmic stability to derive generalization error bounds, characterizing the optimal synthetic-to-real data ratio that minimizes expected test error as a function of the Wasserstein distance between the real and synthetic distributions. We motivate our framework in the setting of kernel ridge regression with mixed data, offering a detailed analysis that may be of independent interest. Our theory predicts the existence of an optimal ratio, leading to a U-shaped behavior of test error with respect to the proportion of synthetic data. Empirically, we validate this prediction on CIFAR-10 and a clinical brain MRI dataset. Our theory extends to the important scenario of domain adaptation, showing that carefully blending synthetic target data with limited source data can mitigate domain shift and enhance generalization. We conclude with practical guidance for applying our results to both in-domain and out-of-domain scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。