arXiv:2608.11426cs.CL2026-08

大模型输出同质化根源在预训练阶段,对齐只是放大了问题。

Is Convergence Inevitable? Tracing Output Homogeneity Back to Base Models

论文配图:Is Convergence Inevitable? Tracing Output Homogeneity Back to Base Models
图 1 · 摘自论文原文
  • 通过控制实验发现,指令微调仅会放大而非引入输出同质性。
  • 基础模型在提示下即出现语义收敛,无需对齐过程。
  • 提示工程本身就能引发类似对齐的同质化现象,适合关注生成多样性研究者。

语言模型内容缺乏多样性通常归因于对齐过程,但其根源何时何地出现尚不明确。本文认为,输出同质性可能在预训练阶段已形成,仅在对齐过程中被揭示或放大。具体而言,我们发现在首个对齐阶段——指令微调(SFT)中即观察到语义收敛,表明同质性可能早已存在于对齐前的基础模型中。通过控制SFT实验,考察训练数据对特定输入/输出对的影响,发现同质性可被揭示和放大,但不会由SFT数据引入,支持其作为催化剂而非原因的角色。进一步测试发现,基础模型在仅使用提示时即出现类似指令对齐的语义坍缩,证明同质性可在无对齐情况下产生。综上,语义收敛可能源于语言模型训练目标的固有特性,单纯依赖对齐后干预难以有效缓解。

原文摘要 · Abstract (English)

The lack of diversity in LM content is widely attributed to the alignment process, but how and where exactly in the pipeline this collapse begins is unknown. We argue that output homogeneity is likely learned during the pretraining phase, and only revealed or magnified during the alignment process. Specifically, we find that semantic convergence is observed from the first alignment stage--the instruction-tuning phase (SFT)--suggesting that homogeneity might already exist in the pre-alignment model. To investigate this, we conduct controlled SFT experiments examining how training data influences output convergence on specific input/output pairs. We find that convergence can be revealed and amplified, but not introduced by the SFT data, supporting its role as a catalyst rather than a cause. To further test whether homogeneity originates before alignment, we measure convergence in base models. We find that instruct-like collapse can be induced through prompting alone, even without alignment. Taken together, our results suggest that semantic convergence may arise naturally from the objectives underlying LM training, making it difficult to mitigate through post-alignment interventions alone.

语言模型生成多样性对齐偏差

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。