arXiv:2503.01832cs.CLcs.LG2025-03被引 2

发现旋转位置编码中隐藏的显著特征,揭示其在大模型中的稳定出现规律。

Rotary Offset Features in Large Language Models

  • 分析旋转编码下查询与键的特征模式,提出旋转偏移特征概念。
  • 发现该特征在多层、多头及不同架构模型中持续出现且激活值显著。
  • 推导出产生该特征的频率与角度阈值,实证验证跨模型一致性。

基于Transformer的大语言模型依赖位置编码为注意力机制提供序列位置信息。旋转位置编码(RoPE)通过旋转查询和键来编码相对位置,已成为现代大模型的主流方案。本文研究使用旋转嵌入时查询与键中涌现的特征与模式,提出旋转偏移特征的概念。分析显示,这些特征频繁表现出显著激活,常被误判为异常值,并在不同层、注意力头和模型架构中稳定出现。我们推导出预测哪些旋转频率会产生旋转偏移特征,以及此类特征对应的查询-键对最小夹角的理论边界,并在多种尺寸与架构的模型上实证验证了预测结果。

原文摘要 · Abstract (English)

Transformer-based Large Language Models (LLMs) rely on positional encodings to provide sequence position information to their attention mechanism. Rotary Positional Encodings (RoPE), which encode relative position by rotating queries and keys, have become widely used in modern LLMs. We study the features and patterns that emerge in queries and keys when using rotary embeddings and introduce the concept of rotary offset features. Our analysis reveals that these features, which frequently exhibit large activations and are often interpreted as outliers, arise consistently across layers, attention heads, and model architectures. We derive bounds predicting which rotary frequencies give rise to rotary offset features and the minimum angle between the query-key pairs for these features. We verify our predictions empirically across models of different sizes and architectures.

位置编码旋转编码大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。