提出加权稳定框架,解决生成模型递归训练中的性能退化问题。
Recursive Learning Without Collapse: A Weighting-Based Stabilization Framework
- 通过混合真实与合成数据,设计加权训练策略提升稳定性。
- 理论证明最优权重随合成数据比例变化呈统一规律,可趋近黄金分割倒数。
- 在模拟与真实表格数据上验证有效,适用于递归生成场景。
近期研究发现,递归生成模型训练中存在模型坍缩现象,即在前序模型生成的数据上训练时性能严重退化。本文提出一种新框架,将生成模型迭代训练于新收集的真实数据与上一轮生成的合成数据组合之上。为优化真实与合成数据的融合策略,评估了多种加权训练方案在高斯分布估计、广义线性模型及非参数估计等场景下的表现。理论上刻画了合成数据混合比例与加权方案对最终模型性能的影响。关键发现:在不同设定下,合成数据不同比例对应的最优加权方案渐近遵循统一表达式,揭示了利用合成数据与模型性能之间的根本权衡。某些情况下,真实数据的最优权重对应黄金比例的倒数。最后,在大量模拟数据集和一个真实表格数据集上验证了理论结果。
原文摘要 · Abstract (English)
Recent studies identified an intriguing phenomenon in recursive generative model training known as model collapse, where models trained on data generated by previous models exhibit severe performance degradation. Addressing this issue and developing more effective training strategies have become central challenges in generative model research. In this paper, we investigate this phenomenon within a novel framework, where generative models are iteratively trained on a combination of newly collected real data and synthetic data from the previous training step. To develop an optimal training strategy for integrating real and synthetic data, we evaluate the performance of a weighted training scheme in various scenarios, including Gaussian distribution estimation, generalized linear models, and nonparametric estimation. We theoretically characterize the impact of the mixing proportion and weighting scheme of synthetic data on the final model's performance. Our key finding is that, across different settings, the optimal weighting scheme under different proportions of synthetic data asymptotically follows a unified expression, revealing a fundamental trade-off between leveraging synthetic data and model performance. In some cases, the optimal weight assigned to real data corresponds to the reciprocal of the golden ratio. Finally, we validate our theoretical results on extensive simulated datasets and a real tabular dataset.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。