AI生成内容泛滥正导致模型退化,研究预测其临界点。
Future of AI Models: A Computational perspective on Model collapse
- 用变换器嵌入分析维基百科语义相似度,追踪语言多样性变化
- 2024年后相似度指数飙升,暗示模型坍塌已进入加速阶段
- 适合关注AI伦理、数据质量与长期模型可持续性的研究者
人工智能,尤其是大语言模型(LLMs),已深刻改变软件工程、新闻、创意写作、学术和媒体等领域。如Stable Diffusion等扩散模型可从文本生成高质量图像与视频。证据显示:74.2%的新发布网页含AI生成内容(Ryan Law 2025),30-40%活跃网络语料为合成内容(Spennemann 2025;arXiv:2504.08755),52%美国成年人使用LLMs进行写作、编程或研究(Staff 2025),审计发现18%金融投诉与24%新闻稿涉及AI(Liang et al. 2025)。现有神经架构依赖大规模、多样化的真人创作数据集。随着合成内容主导,递归训练正威胁语言与语义多样性,引发模型坍塌(Shumailov et al. 2024;arXiv:2307.15043;Dohmatob et al. 2024;arXiv:2402.07712)。本研究通过分析2013–2025年英语维基百科(过滤Common Crawl)的年份级语义相似度,结合变换器嵌入与余弦相似度度量,量化并预测坍塌发生时间。结果显示:在公共LLM采用前,相似度稳步上升,可能源于早期RNN/LSTM翻译与文本标准化流程,但幅度有限;波动反映不可消除的语言多样性、年度语料规模差异、有限采样误差。而公共采用后,相似度呈指数增长。结果提供了数据驱动的模型数据丰富性与泛化能力受威胁的临界点预估。
原文摘要 · Abstract (English)
Artificial Intelligence, especially Large Language Models (LLMs), has transformed domains such as software engineering, journalism, creative writing, academia, and media (Naveed et al. 2025; arXiv:2307.06435). Diffusion models like Stable Diffusion generate high-quality images and videos from text. Evidence shows rapid expansion: 74.2% of newly published webpages now contain AI-generated material (Ryan Law 2025), 30-40% of the active web corpus is synthetic (Spennemann 2025; arXiv:2504.08755), 52% of U.S. adults use LLMs for writing, coding, or research (Staff 2025), and audits find AI involvement in 18% of financial complaints and 24% of press releases (Liang et al. 2025). The underlying neural architectures, including Transformers (Vaswani et al. 2023; arXiv:1706.03762), RNNs, LSTMs, GANs, and diffusion networks, depend on large, diverse, human-authored datasets (Shi & Iyengar 2019). As synthetic content dominates, recursive training risks eroding linguistic and semantic diversity, producing Model Collapse (Shumailov et al. 2024; arXiv:2307.15043; Dohmatob et al. 2024; arXiv:2402.07712). This study quantifies and forecasts collapse onset by examining year-wise semantic similarity in English-language Wikipedia (filtered Common Crawl) from 2013 to 2025 using Transformer embeddings and cosine similarity metrics. Results reveal a steady rise in similarity before public LLM adoption, likely driven by early RNN/LSTM translation and text-normalization pipelines, though modest due to a smaller scale. Observed fluctuations reflect irreducible linguistic diversity, variable corpus size across years, finite sampling error, and an exponential rise in similarity after the public adoption of LLM models. These findings provide a data-driven estimate of when recursive AI contamination may significantly threaten data richness and model generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。