arXiv:2601.04098cs.CLcs.AI2026-01被引 2

揭示了语言模型在短文本中逐层的定位偏好机制。

Layer-wise Positional Bias in Short-Context Language Modeling

  • 提出滑动窗口下的层传导分析框架,分离内部结构与任务干扰。
  • 发现模型逐层递增的近期偏好,内容词更倾向首部位置。
  • 函数词受近期偏倚影响更大,适用于模型可解释性研究。

Transformer语言模型在输入中存在系统性位置偏好,即对特定位置的词有非语义相关的倾向,称为位置偏倚。以往研究通过长上下文任务性能下降或注意力分析来刻画该现象,但缺乏对各层内部行为的逐层测量。本文引入基于滑动窗口的层传导框架,应用于短上下文的下一个词预测任务,以隔离模型内部行为与任务及上下文窗口压力的影响。得到的逐层位置重要性分布稳定,且在不同文本和词汇打乱下保持一致,证实其反映模型内部结构。分析表明,近期偏倚随网络深度单调上升,而首因偏倚较弱并逐渐减弱;此外,位置偏倚并非均匀分布:功能词表现出更强的近期偏倚,而内容词则显示出更高的首因偏倚。

原文摘要 · Abstract (English)

Transformer language models systematically prefer tokens at specific input positions regardless of semantic relevance---a phenomenon known as positional bias. Prior work characterizes this bias in model behavior through performance drops in long-context tasks or in model architecture through attention-based analyses. However, it remains unmeasured how input positions actually drive predictions layer by layer. We introduce a layer conductance framework within a sliding-window design, applied to short-context next-word prediction to isolate model-internal behavior from task and context-window pressure. The resulting layer-wise positional importance profiles are stable across diverse texts and lexical scrambling, confirming they reflect model-internal structure. Characterizing how these profiles evolve across depth, we find recency bias increases monotonically while primacy bias is subtle and diminishes. We also find that this positional bias is not uniform across word types: function words exhibit higher recency bias while content words show higher primacy bias.

位置偏倚可解释性语言模型逐层分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。