用动态门控和旋转嵌入提升对称幂变压器的长序列记忆能力
Conformal Transformations for Symmetric Power Transformers
- 引入数据依赖的门控机制释放递归状态容量
- 在LongCrawl64上实现训练与推理长度扩展时的稳定性能
- 适合需要长序列建模的高效注意力研究者
基于线性注意力的Transformer相比Softmax Transformer具有显著计算优势,但常伴随性能下降。对称幂(sympow)Transformer通过利用对称张量嵌入,在部分缓解性能差距方面取得进展,达到与Softmax Transformer相当的水平。然而,其递归状态容量有限,导致在训练或评估上下文长度扩展时性能退化。为此,本文提出共形-对称幂(conformal-sympow)Transformer,通过数据依赖的乘法门控动态释放容量,并采用数据依赖的旋转嵌入自适应存储信息。初步实验在LongCrawl64数据集上表明,conformal-sympow克服了sympow的局限性,在扩展的训练与评估上下文长度下仍保持稳健表现。
原文摘要 · Abstract (English)
Transformers with linear attention offer significant computational advantages over softmax-based transformers but often suffer from degraded performance. The symmetric power (sympow) transformer, a particular type of linear transformer, addresses some of this performance gap by leveraging symmetric tensor embeddings, achieving comparable performance to softmax transformers. However, the finite capacity of the recurrent state in sympow transformers limits their ability to retain information, leading to performance degradation when scaling the training or evaluation context length. To address this issue, we propose the conformal-sympow transformer, which dynamically frees up capacity using data-dependent multiplicative gating and adaptively stores information using data-dependent rotary embeddings. Preliminary experiments on the LongCrawl64 dataset demonstrate that conformal-sympow overcomes the limitations of sympow transformers, achieving robust performance across scaled training and evaluation contexts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。