arXiv:2509.10509cs.LGcs.AI2025-09

筛选自生成内容可让大模型越训越强,打破自我迭代衰减魔咒。

The Anti-Ouroboros Effect: Emergent Resilience in Large Language Models from Recursive Selective Feedback

  • 用质量筛选机制对自生成数据进行递归训练
  • 2B模型在摘要任务上性能提升6.6%,优于无过滤和随机过滤
  • 适合关注AI安全与自进化系统的研究者

大语言模型的递归训练稳定性是人工智能安全的核心问题。主流理论预测模型会因自生成输出训练而发生退化(模型崩溃)。我们提出一种选择性反馈机制,挑战这一观点:实验表明,该机制不仅延缓退化,更逆转了趋势,在复杂摘要任务中使Gemma 2B模型性能显著提升。我们称此现象为“反奥罗伯洛斯效应”。与简单分类器中验证的退化循环对比,凸显出高维模型的独特动态。研究发现,仅通过简单筛选压力,系统韧性可成为大模型的涌现特性,为构建更安全、鲁棒的AI系统提供可扩展原则。五代训练中,质量筛选组ROUGE-L F1提升6.6%,未过滤对照组下降3.5%,随机筛选组下降4.2%。

原文摘要 · Abstract (English)

The stability of recursively trained large language models (LLMs) is a foundational problem for AI safety. Prevailing theory predicts model collapse, a progressive degradation when models are trained on their own output. We challenge this narrative by introducing a selective feedback mechanism. Contrary to expectation, instead of merely slowing decay, our experiments provide strong evidence that this pressure reverses it, inducing a statistically significant performance improvement in a Gemma 2B model on a complex summarization task. We name this phenomenon the Anti-Ouroboros Effect. We contrast this with a foundational experiment using a simple classifier, where the theoretical degenerative loop was validated, highlighting the unique dynamics of high-dimensional models. Our findings establish that systemic resilience can be an emergent property of LLMs under simple selection pressure, suggesting a powerful and scalable principle for developing safer and more robust AI systems. Across five generations, a quality-filtered condition improved by 6.6% in ROUGE-L F1 score, whereas an unfiltered control degraded by 3.5% and a random-filter control degraded by 4.2%

大模型训练自进化AI安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。