训练目标比模型规模更影响语言风格,当前对齐机制正系统性扭曲语言分布。
From Context Shift to Stylistic Collapse: Why Training Objectives Matter More Than Scale

- 用24个语言探针分析17个模型,发现指令微调使语言熵放大1949%-16853%
- 复杂标点被压制至基线频率的3.2%-23.2%,且强化学习无改善作用
- 强正则化(lambda=5.0)可显著恢复语言多样性,优于顶尖大模型
现代语言模型中,语言特征不再作为风格痕迹,而是概率质量的探测器,其分配受训练对齐目标主导。我们分析了17个模型(参数量410M-100B+)在24个语言动机探针下的表现,发现指令微调系统性地重塑语言特征,在话语与结构维度上导致极端语言重分布(平均放大1,949-16,853%,峰值达5,181-209,675%),同时选择性抑制复杂标点至基线频率的3.2-23.2%。该现象在强化学习人类反馈(RLHF)下未加剧,匹配基线与指令微调模型对差异不显著(p > 0.25)。弱干预(lambda=1.0)使崩溃恶化240%,而强控制(lambda=5.0)实现40.5%性能提升,虽仅具200-1000倍规模劣势,仍超越前沿模型96.7-98.2%。λ=5.0还带来15%更高的distinct-4、27%更高的词汇多样性、78%更低重复率,表明对齐需足够控制强度而非仅分布平滑。研究揭示当前对齐流程存在结构性缺陷:偏好优化重塑语言分布,但标准质量指标无法捕捉,仅可通过分布探针检测,对AI识别、训练数据污染及长期语言演化有深远影响。
原文摘要 · Abstract (English)
In modern LLMs, linguistic features function not as stylistic artifacts but as probes of probability mass, allocated under training alignment objectives. Language models trained with contemporary pipelines exhibit severe reshaping of linguistic features, leading to extreme language re-distribution. While previous stylometric analyses explored linguistic differences between AI-generated and human texts, we focus on the reshaping plaguing the LLM training pipeline itself. We analyze 17 models (410M-100B+ parameters) across 24 linguistically-motivated probes, documenting that instruction-tuned systems systematically collapse language entropy along discourse and structural dimensions (mean amplification: 1,949-16,853%, peaks: 5,181-209,675%), while selectively suppressing complex punctuation to 3.2-23.2% of baseline frequencies. These effects do not worsen under RLHF, as divergence patterns are statistically indistinguishable (p > 0.25) across matched base and instruction-tuned model pairs. Weak intervention (lambda=1.0) exacerbates collapse by 240%, while strong control (lambda=5.0) achieves 40.5% improvement and outperforms frontier models by 96.7-98.2% despite 200-1000x scale disadvantage. Additionally, lambda=5.0 delivers 15% higher distinct-4, 27% higher vocabulary diversity, and 78% lower repetition than moderate regularization, establishing that alignment requires sufficient control strength, not merely distributional smoothing. Our findings underscore how modern LLMs reallocate stylistic probability mass, despite RLHF and scale. More broadly, our work reveals a structural limitation of current alignment pipelines: preference optimization reshapes language distributions invisible to standard quality metrics yet detectable through distributional probes, with implications for AI detection, training data contamination, and long-term linguistic evolution.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。