arXiv:2503.03150cs.LGcs.AI2025-03被引 33

模型崩溃并非普遍威胁,实为术语混乱导致的误判。

Position: Model Collapse Does Not Mean What You Think

  • 澄清八种不同定义的模型崩溃,指出术语不统一问题
  • 实证分析显示多数崩溃场景在真实条件下可避免
  • 提醒关注更紧迫的现实风险,而非夸大技术威胁

AI生成内容泛滥引发对'模型崩溃'的担忧,即未来生成模型因训练于早期模型生成的合成数据而性能下降。行业领袖、顶级期刊与大众媒体均预言其将带来灾难性社会后果。本文认为这一主流叙事严重误解了科学证据。研究发现,'模型崩溃'实际包含八种互不一致的定义,术语混乱阻碍了深入理解。通过设定合理现实条件评估文献方法,我们发现多数预测依赖脱离实际的假设,部分典型崩溃情景实际上可轻易规避。因此,模型崩溃被简化为单一威胁,而真正可能的风险却未受足够重视。

原文摘要 · Abstract (English)

The proliferation of AI-generated content online has fueled concerns over \emph{model collapse}, a degradation in future generative models' performance when trained on synthetic data generated by earlier models. Industry leaders, premier research journals and popular science publications alike have prophesied catastrophic societal consequences stemming from model collapse. In this position piece, we contend this widespread narrative fundamentally misunderstands the scientific evidence. We highlight that research on model collapse actually encompasses eight distinct and at times conflicting definitions of model collapse, and argue that inconsistent terminology within and between papers has hindered building a comprehensive understanding of model collapse. To assess how significantly different interpretations of model collapse threaten future generative models, we posit what we believe are realistic conditions for studying model collapse and then conduct a rigorous assessment of the literature's methodologies through this lens. While we leave room for reasonable disagreement, our analysis of research studies, weighted by how faithfully each study matches real-world conditions, leads us to conclude that certain predicted claims of model collapse rely on assumptions and conditions that poorly match real-world conditions, and in fact several prominent collapse scenarios are readily avoidable. Altogether, this position paper argues that model collapse has been warped from a nuanced multifaceted consideration into an oversimplified threat, and that the evidence suggests specific harms more likely under society's current trajectory have received disproportionately less attention.

模型崩溃人工智能伦理术语辨析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。