研究生成模型自我迭代时的崩溃问题,给出防止崩溃的最佳数据混合比例。
Preventing Model Collapse Under Overparametrization: Optimal Mixing Ratios for Interpolation Learning and Ridge Regression
- 通过迭代混合真实与合成标签,分析线性回归中模型性能。
- 最优真实数据比例收敛到黄金分割倒数,且至少需占一半。
- 适用于自训练、数据分布变化等场景,为生成模型提供理论指导。
模型崩溃指生成模型在反复训练自身合成输出后性能下降。本文研究在过参数化线性回归下,每轮迭代混合新真实标签与前一轮模型生成的合成标签时的崩溃现象。推导了最小ℓ₂范数插值与岭回归的精确泛化误差公式。分析表明,长期预测误差最小时的最优混合权重具有关键性质:在最小ℓ₂范数插值情形下,最优真实数据比例对广泛协变量分布均收敛至黄金比例的倒数(约0.618),此前仅知于普通最小二乘且限于低维情况;对于岭回归,进一步分析随机效应模型与尖峰协方差模型,揭示谱几何如何决定最优权重。在所有情形(包括各向同性特征)下,最优混合比均不低于1/2,表明必须优先使用真实数据。还考察三种扩展设置:(i) 真实数据固定但无新标签;(ii) 特征随轮次变化但有新标签;(iii) 特征随时间变化,仅部分特征获得新标签。在这些设定中,明确刻画了模型崩溃不可避免的条件及合成数据提升学习的边界。通过大量模拟验证理论结果。
原文摘要 · Abstract (English)
Model collapse occurs when generative models degrade after repeatedly training on their own synthetic outputs. We study this effect in overparameterized linear regression in a setting where each iteration mixes fresh real labels with synthetic labels drawn from the model fitted in the previous iteration. We derive precise generalization error formulae for minimum-$\ell_2$-norm interpolation and ridge regression under this iterative scheme. Our analysis reveals intriguing properties of the optimal mixing weight that minimizes long-term prediction error and provably prevents model collapse. For instance, in the case of min-$\ell_2$-norm interpolation, we establish that the optimal real-data proportion converges to the reciprocal of the golden ratio for fairly general classes of covariate distributions. Previously, this property was known only for ordinary least squares, and additionally in low dimensions. For ridge regression, we further analyze two popular model classes -- the random-effects model and the spiked covariance model -- demonstrating how spectral geometry governs optimal weighting. In both cases, as well as for isotropic features, we uncover that the optimal mixing ratio should be at least one-half, reflecting the necessity of favoring real-data over synthetic. We study three additional settings: (i) where real data is fixed and fresh labels are not obtained at each iteration, (ii) where covariates vary across iterations but fresh real labels are available each time, and (iii) where covariates vary with time but only a fraction of them receive fresh real labels at each iteration. Across these diverse settings, we characterize when model collapse is inevitable and when synthetic data improves learning. We validate our theoretical results with extensive simulations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。