研究大模型生成文本如何自我强化,揭示信息衰减与筛选机制的双重影响。
Drift and selection in LLM text ecosystems

- 用变阶n-gram模型构建递归学习框架,分离出文本演化中的漂移与选择机制。
- 无限语料下漂移导致罕见表达消失,而筛选可维持深层结构并限制偏离浅层均衡。
- 适用于关注大模型训练数据质量、设计可持续文本生态的研究者。
公共文本记录——当前人与人工智能系统共同学习的材料——正越来越多地被其自身输出所塑造。生成文本进入公共记录,后续智能体从中学习,循环往复。本文提出一个可精确求解的数学框架,基于变阶n-gram代理,分离出作用于公共语料库的两种力量:第一是漂移——无过滤复用逐步消除稀有形式,在无限语料极限下可精确刻画稳定分布;第二是选择——发布、排序与验证对进入记录的内容进行筛选,结果取决于被选中的内容。当发布仅反映统计现状时,语料库趋于浅层状态,进一步前瞻不再带来收益;当发布具有规范性——奖励质量、正确性或新颖性——则深层结构得以保留,并建立其与浅层均衡之间偏移的最优上界。该框架因此揭示了递归发布在压缩公共文本与维持丰富结构之间的权衡,对人工智能训练语料的设计具有重要意义。
原文摘要 · Abstract (English)
The public text record -- the material from which both people and AI systems now learn -- is increasingly shaped by its own outputs. Generated text enters the public record, later agents learn from it, and the cycle repeats. Here we develop an exactly solvable mathematical framework for this recursive process, based on variable-order $n$-gram agents, and separate two forces acting on the public corpus. The first is drift: unfiltered reuse progressively removes rare forms, and in the infinite-corpus limit we characterise the stable distributions exactly. The second is selection: publication, ranking and verification filter what enters the record, and the outcome depends on what is selected. When publication merely reflects the statistical status quo, the corpus converges to a shallow state in which further lookahead brings no benefit. When publication is normative -- rewarding quality, correctness or novelty -- deeper structure persists, and we establish an optimal upper bound on the resulting divergence from shallow equilibria. The framework therefore identifies when recursive publication compresses public text and when selective filtering sustains richer structure, with implications for the design of AI training corpora.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。