提出Wonderful Matrices架构,融合序列与状态变换提升模型效率与效果。
Wonderful Matrices: Combining for a More Efficient and Effective Foundation Model Architecture
- 结合旋转位置编码与状态空间对偶,统一位置信息表示。
- 动态掩码注意力在多查询回忆任务中达100%准确率,性能超基线150%以上。
- 跨域专家混合加速专家检索,支持千级专家时提速8-10倍,适合高效大模型设计。
为提升基础模型的效率与效果,本文提出融合序列变换与状态变换的新范式。首先,证明旋转位置编码在状态空间对偶算法中的有效性,使混合二次因果自注意力与状态空间对偶的困惑度降低4%以上,确保序列变换中位置编码的一致性。其次,提出动态掩码注意力,在更具挑战性的多查询关联回忆任务中保持100%准确率,相较二次因果自注意力与状态空间对偶提升超过150%,实现关键信息的选择性过滤。第三,设计跨域专家混合机制,使包含超过1024个专家的专家检索速度比传统混合专家快8至10倍,实现状态变换下的快速专家调用。最终,整合上述矩阵算法形成名为Wonderful Matrices的基础模型架构,可作为主流模型架构的有力竞争者。
原文摘要 · Abstract (English)
In order to make the foundation model more efficient and effective, our idea is combining sequence transformation and state transformation. First, we prove the availability of rotary position embedding in the state space duality algorithm, which reduces the perplexity of the hybrid quadratic causal self-attention and state space duality by more than 4%, to ensure that the combining sequence transformation unifies position encoding. Second, we propose dynamic mask attention, which maintains 100% accuracy in the more challenging multi-query associative recall task, improving by more than 150% compared to quadratic causal self-attention and state space duality, to ensure that the combining sequence transformation selectively filters relevant information. Third, we design cross domain mixture of experts, which makes the computational speed of expert retrieval with more than 1024 experts 8 to 10 times faster than the mixture of experts, to ensure that the combining state transformation quickly retrieval mixture. Finally, we summarize these matrix algorithms that can form the foundation model: Wonderful Matrices, which can be a competitor to popular model architectures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。