自训练让语言表面更复杂,深层语法却在退化。
Self-Training Doesn't Flatten Language -- It Restructures It: Surface Markers Amplify While Deep Syntax Dies

- 按句法深度预测语言特征衰减,深度越深越易消失
- 生成文本中连接词增多,疑问、被动等结构大幅减少
- 适合关注模型幻觉与数据质量的研究者阅读
对语言模型自训练过程的多次迭代普遍被描述为‘扁平化’:多样性下降,分布收窄,文本趋于自我复制。我们提供证据表明这一描述不完整。在五个模型(GPT-2 124M、Pythia-410M、Pythia-1.4B、OPT-1.3B、Pythia-2.8B)上进行十一轮自训练,语言并非均匀扁平化,而是结构性重组。表面标记(如连接词、缓和语、破折号)增加,而中层与深层句法结构(如疑问句、插入语、被动语态、虚拟语气)显著坍塌。我们提出结构深度假说(SDH):语言特征每代衰减速率主要由其结构深度(嵌套句法依赖数)决定,其次才是初始输出频率。整合五模型共17项特征(覆盖三个架构族,总计85个样本),总斯皮尔曼相关系数rho=0.540(p < 10^{-6};聚类自助95%置信区间[0.434, 0.634]),频率则弱得多(rho=0.225)。人类文本微调对照组rho=0.039(p=0.88),确认该梯度为自训练特有。此外发现‘表层复杂性悖论’:尽管深层从句结构消亡,总体复杂性指标(依存树深度、类型-词频比TTR、词长)均上升,这对训练数据筛选与大模型文本检测有直接意义。
原文摘要 · Abstract (English)
Successive self-training on a language model's own outputs is widely characterized as a process of flattening: diversity drops, distributions narrow, and the text becomes "more like itself." We provide evidence that this characterization is incomplete. Across eleven generations of self-training on five models (GPT-2 124M, Pythia-410M, Pythia-1.4B, OPT-1.3B, Pythia-2.8B), language is not flattened uniformly -- it is restructured. Surface markers (discourse connectives, hedges, em-dashes) rise, while mid- and deep-syntactic structures (questions, parentheticals, passives, subjunctives) collapse. We formalize this asymmetric collapse as the Structural Depth Hypothesis (SDH): the per-generation decay rate of a linguistic feature is predicted primarily by its structural depth -- the number of nested syntactic dependencies it requires -- and only secondarily by its generation-zero output frequency. Pooling 17-feature panels from five models spanning three architecture families (N=85), the pooled Spearman correlation is rho=0.540 (p < 10^{-6}; cluster-bootstrap 95% CI [0.434, 0.634]), while frequency is a substantially weaker predictor (rho=0.225). A matched human-text fine-tuning control yields rho=0.039 (p=0.88), confirming the gradient is self-training-specific. We further document a Superficial Complexity Paradox: aggregate complexity proxies (dep-tree depth, TTR, word length) all rise as the underlying clause structure dies, with direct implications for training-data curation and LLM-text detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。