arXiv:2505.21677cs.LGcs.AI2025-05被引 3

研究AI模型互训时生成内容的演化与影响

What happens when generative AI models train recursively on each others' outputs?

  • 让不同AI模型互相用对方生成的内容训练
  • 能引入新概念但导致性能趋同
  • 对依赖AI的行业有重要警示意义

互联网是生成式AI(genAI)模型的共同训练数据源,但正越来越多地被AI生成内容填充。这种双重性引发了一种可能性:未来的genAI模型可能在其他模型生成的输出上进行训练。以往研究主要关注模型自训其生成内容的后果,但较少探讨模型摄入其他模型产出内容的影响。鉴于社会对genAI工具的依赖日益加深,理解此类数据驱动的模型交互至关重要。本文提供了实证证据,揭示了实际中此类交互如何展开;构建了该交互训练过程的理论模型,并通过实验验证了该理论。研究发现,数据驱动的交互可使模型接触到原始训练数据中遗漏的新概念,从而带来收益;但也可能导致它们在共享任务上的表现趋于同质化。

原文摘要 · Abstract (English)

The internet serves as a common source of training data for generative AI (genAI) models but is increasingly populated with AI-generated content. This duality raises the possibility that future genAI models may be trained on other models' generated outputs. Prior work has studied consequences of models training on their own generated outputs, but limited work has considered what happens if models ingest content produced by other models. Given society's increasing dependence on genAI tools, understanding such data-mediated model interactions is critical. This work provides empirical evidence for how data-mediated interactions might unfold in practice, develops a theoretical model for this interactive training process, and experimentally validates the theory. We find that data-mediated interactions can benefit models by exposing them to novel concepts perhaps missed in original training data, but also can homogenize their performance on shared tasks.

生成式AI模型互训数据污染

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。