发现大模型中位置偏差会加剧令牌表征趋同现象。
Token Homogenization under Positional Bias
- 通过逐层相似性分析,发现令牌表征随深度增加趋于一致。
- 极端位置的令牌受注意力机制影响,表征失真更严重。
- 揭示了位置偏差如何放大模型内部表征的同质化问题。
本文研究了大型语言模型中的令牌同质化现象——即在变压器网络中,令牌表征随层加深逐渐趋同,以及其与位置偏差的关系。通过逐层相似性分析和受控实验,我们验证了该现象的存在,并发现当令牌受极端位置偏向时,这种同质化效应被显著放大。结果表明,令牌表征的多样性在深层网络中系统性下降,且这一过程与位置注意力机制密切相关。
原文摘要 · Abstract (English)
This paper investigates token homogenization - the convergence of token representations toward uniformity across transformer layers and its relationship to positional bias in large language models. We empirically examine whether homogenization occurs and how positional bias amplifies this effect. Through layer-wise similarity analysis and controlled experiments, we demonstrate that tokens systematically lose distinctiveness during processing, particularly when biased toward extremal positions. Our findings confirm both the existence of homogenization and its dependence on positional attention mechanisms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。