arXiv:2504.03814cs.LGcs.AI2025-04EMNLP被引 3

研究互联网数据特性如何影响大模型生成内容的分布偏移

Recursive Training Loops in LLMs: How training data properties modulate distribution shift in generated data?

  • 通过操控数据集属性,发现词汇多样性加剧分布偏移
  • 语义多样性和数据质量可缓解偏移,且影响具有领域隔离性
  • 揭示了不同网络领域可能经历不同类型的分布偏移

大型语言模型(LLMs)越来越多用于在线内容生成,形成训练数据与生成内容间的反馈循环。此类循环可能导致分布偏移——模型无法反映人类数据的真实分布(也称模型坍塌)。然而,人类数据属性如何影响这种偏移仍不明确。本文首次实证考察数据属性对递归训练结果的影响。我们首先确认使用不同人类数据集会导致不同程度的分布偏移。通过系统性操纵数据集属性并结合回归分析,识别出若干预测偏移程度的关键属性:词汇多样性会放大偏移,而语义多样性和数据质量则能缓解偏移。此外,这些影响具有高度模块化特征:某一互联网领域的数据对另一领域生成内容影响甚微。最后,在政治偏见实验中发现,人类数据属性决定了初始偏见是被放大还是被削弱。总体而言,研究揭示了一种新图景:互联网不同部分可能经历不同类型的分布偏移。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly used in the creation of online content, creating feedback loops as subsequent generations of models will be trained on this synthetic data. Such loops were shown to lead to distribution shifts - models misrepresenting the true underlying distributions of human data (also called model collapse). However, how human data properties affect such shifts remains poorly understood. In this paper, we provide the first empirical examination of the effect of such properties on the outcome of recursive training. We first confirm that using different human datasets leads to distribution shifts of different magnitudes. Through exhaustive manipulation of dataset properties combined with regression analyses, we then identify a set of properties predicting distribution shift magnitudes. Lexical diversity is found to amplify these shifts, while semantic diversity and data quality mitigate them. Furthermore, we find that these influences are highly modular: data scrapped from a given internet domain has little influence on the content generated for another domain. Finally, experiments on political bias reveal that human data properties affect whether the initial bias will be amplified or reduced. Overall, our results portray a novel view, where different parts of internet may undergo different types of distribution shift.

大模型分布偏移数据质量反馈循环

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。