arXiv:2605.11206cs.CL2026-05

指令主要影响语言模型生成,而非理解。

Instructions Shape Production of Language, not Processing

论文配图:Instructions Shape Production of Language, not Processing
图 1 · 摘自论文原文
  • 通过分层探测输入与输出信息,发现指令显著影响生成阶段。
  • 输出中的任务信息随指令变化明显,且与行为强相关。
  • 该机制在大模型和指令微调后更显著,适合研究模型生成逻辑者。

指令触发语言模型的生成中心机制。基于区分语言处理与生成的认知视角,我们通过五种二分类判断任务,在层间探测任务特定信息,揭示了两个阶段间的不对称性。具体而言,测量指令标记如何影响样本标记(待评估输入)的处理和输出标记的生成。在不同提示方式下,样本标记中的任务信息基本稳定,与行为关联微弱;而输出标记中的相同信息变化显著,与行为高度相关。基于注意力的干预实验确认此模式具有因果性:阻断指令流向所有后续标记会同时降低行为表现和输出信息;仅阻断指令流向样本标记则对两者影响甚微。该不对称性在不同模型家族和任务中均成立,且随模型规模和指令微调程度增强,均更显著影响生成阶段。研究提示,理解模型能力需结合内部状态与行为表现,并按标记位置分解内部视角,以区分输入处理与输出生成。

原文摘要 · Abstract (English)

Instructions trigger a production-centered mechanism in language models. Through a cognitively inspired lens that separates language processing and production, we reveal this mechanism as an asymmetry between the two stages by probing task-specific information layer-wise across five binary judgment tasks. Specifically, we measure how instruction tokens shape information both when sample tokens, the input under evaluation, are processed and when output tokens are produced. Across prompting variations, task-specific information in sample tokens remains largely stable and correlates only weakly with behavior, whereas the same information in output tokens varies substantially and correlates strongly with behavior. Attention-based interventions confirm this pattern causally: blocking instruction flow to all subsequent tokens reduces both behavior and information in output tokens, whereas blocking it only to sample tokens has minimal effect on either. The asymmetry generalizes across model families and tasks, and becomes sharper with model scale and instruction-tuning, both of which disproportionately affect the production stage. Our findings suggest that understanding model capabilities requires jointly assessing internals and behavior, while decomposing the internal perspective by token position to distinguish the processing of input tokens from the production of output tokens.

语言模型生成机制指令作用内部机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。