arXiv:2608.21366cs.AIcs.LG2026-08中稿 · and published in P…

梳理生成模型自循环训练中的崩溃现象及应对策略

Reviewing Model Collapse and Countermeasures

论文配图:Reviewing Model Collapse and Countermeasures
图 1 · 摘自论文原文
  • 系统回顾生成模型与合成数据互构导致的崩溃问题
  • 总结不同场景下模型崩溃的典型表现与成因
  • 适合关注AI可信性与数据闭环风险的研究者

得益于海量网络数据,生成式人工智能(GenAI)取得显著进展,推动了多个领域的应用。为应对日益增长的数据需求,从业者开始使用AI生成的数据训练下一代模型。虽然合成数据缓解了数据供给压力,但也引入新问题:模型与数据在自循环中相互依赖,最终导致模型崩溃,引发对生成式AI可信性的担忧。近年来,越来越多研究关注模型崩溃(MC)现象并探索缓解方法,但该问题的系统性综述仍为空白。本文旨在填补这一空白,全面梳理近年相关研究,总结不同应用场景中模型崩溃的表现、成因及应对策略,并指出当前挑战与未来研究方向。

原文摘要 · Abstract (English)

Driven by massive amounts of web-scale data, generative AI (GenAI) has achieved remarkable progress, enabling various applications in diverse sectors. The advances of GenAI have actuated practitioners to use AI-synthesized data for training next-generation AI models. Undeniably, using synthetic data has alleviated the increasing stringent demand for data supply. Unfortunately, it also introduces a new critical issue: in a self-consuming cycle between model and data, the model ultimately collapse, raising more trustworthiness concerns to GenAI. In recent years, increasingly more studies have investigated the phenomenon of model collapse (MC) and explored potential solutions to mitigate it. However, the review of the phenomenon of MC still remains blank. To fill this gap, this paper provides an up-to-date overview of these studies for consolidating and reviewing the progress of MC in different application scenarios and countermeasures for mitigating MC. We also highlight challenges and future research opportunities.

模型崩溃生成模型数据闭环可信AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。