arXiv:2510.04130cs.LGcs.AI2025-10TPAMI

解析位置嵌入对长度泛化的制约与潜力,提出新理论框架与优化方法。

On the Limitations and Capabilities of Position Embeddings for Length Generalization

  • 引入线性表示复杂度(LRC)分析位置嵌入的结构作用
  • 实验证明长度泛化依赖于表示复杂度在尺度间保持不变
  • 提出可伸缩提示与自学习位置嵌入,提升模型泛化能力

在Transformer中,位置嵌入(PEs)显著影响长度泛化(LG)性能,但其根本作用尚不明确。本文理论分析了仅位置线性注意力(POLAs)中的位置嵌入,提出线性表示复杂度(LRC)以刻画位置嵌入实现长度泛化的条件。分析表明,位置嵌入不扩展计算能力,而是结构化跨位置的学习计算。拓展至实际Transformer,提出序列表示复杂度(SRC),并推测长度泛化成立当且仅当SRC在不同尺度下保持不变。通过多种推理任务的实证支持该假设。为增强长度泛化,提出尺度提示(Scale Hint)实现灵活实例缩放,以及基于学习的位置嵌入框架自动捕捉位置关系。本工作提供理论洞见与实用策略,推动Transformer在长度泛化上的改进。

原文摘要 · Abstract (English)

In Transformers, Position Embeddings (PEs) significantly influence Length Generalization (LG) performance, yet their fundamental role remains unclear. In this work, we investigate the limitations and capabilities of PEs in achieving LG. We theoretically analyze PEs in Position-Only Linear Attentions (POLAs), introducing Linear Representation Complexity (LRC) to characterize when PEs enable LG. Our analysis shows that PEs do not expand computational capabilities but structure learned computations across positions. Extending to practical Transformers, we propose Sequential Representation Complexity (SRC) and conjecture that LG is possible if and only if SRC remains invariant across scales. We support this hypothesis with empirical evidence in various reasoning tasks. To enhance LG, we introduce Scale Hint, allowing flexible instance scaling, and a Learning-Based Position Embedding framework that automatically learns positional relations. Our work provides theoretical insights and practical strategies for improving LG in Transformers.

位置嵌入长度泛化Transformer理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。