揭示了模型合成中避免退化的通用规律,关键在于数据增广方式。
Universality of the $π^2/6$ Pathway in Avoiding Model Collapse
- 提出统一框架,证明pi²/6风险上限适用于多种经典统计模型。
- 实验证明增广流程可避免模型退化,而丢弃流程会引发崩溃。
- 适合关注生成模型稳定性与训练流程设计的研究者。
机器学习研究者担忧模型退化现象,即在仅用合成数据迭代训练时模型性能持续下降。此前研究表明,若丢弃真实数据、仅用模型生成的合成数据进行训练,则会出现模型退化;而若持续使用原始真实数据,并加入历史生成的合成数据(增广流程),则能避免退化。针对线性回归,已有理论证明增广流程下测试风险始终不超过原始训练风险的π²/6倍。本文证明该π²/6界限在一大类经典统计模型中具有普适性,揭示了丢弃流程导致退化、增广流程防止退化的根本机制。同时构建了一个可模拟多种训练流程的高斯过程框架,便于评估不同策略的优劣。
原文摘要 · Abstract (English)
Researchers in empirical machine learning recently spotlighted their fears of so-called Model Collapse. They imagined a discard workflow, where an initial generative model is trained with real data, after which the real data are discarded, and subsequently, the model generates synthetic data on which a new model is trained. They came to the conclusion that models degenerate as model-fitting generations proceed. However, other researchers considered an augment workflow, where the original real data continue to be used in each generation of training, augmented by synthetic data from models fit in all earlier generations. Empirical results on canonical datasets and learning procedures confirmed the occurrence of model collapse under the discard workflow and avoidance of model collapse under the augment workflow. Under the augment workflow, theoretical evidence also confirmed avoidance in particular instances; specifically, Gerstgrasser et al. (2024) found that for classical Linear Regression, test risk at any later generation is bounded by a moderate multiple, viz. pi-squared-over-6 of the test risk of training with the original real data alone. Some commentators questioned the generality of theoretical conclusions based on the generative model assumed in Gerstgrasser et al. (2024): could similar conclusions be reached for other task/model pairings? In this work, we demonstrate the universality of the pi-squared-over-6 augment risk bound across a large family of canonical statistical models, offering key insights into exactly why collapse happens under the discard workflow and is avoided under the augment workflow. In the process, we provide a framework that is able to accommodate a large variety of workflows (beyond discard and augment), thereby enabling an experimenter to judge the comparative merits of multiple different workflows by simulating a simple Gaussian process.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。