arXiv:2603.15665cs.AI2026-03被引 1

从语言学出发,揭示注意力机制的本质,提出更高效的QV新范式。

QV May Be Enough: Toward the Essence of Attention in LLMs

  • 基于词性与句法分析,推导出注意力核心是查询与值的匹配
  • 验证QV范式在多种架构中有效,且比传统方法更高效
  • 适合关注模型本质、架构优化的研究者

本文从语言学角度出发,以词性(POS)和句法分析为基础,探究并推导Transformer中查询-键-值(QKV)机制的内在本质。基于此理论基础,我们为当前主流架构(如MQA、GQA、MLA)的性能提供了统一解释框架,并识别其固有权衡及优化方向。提出QV范式,并通过实证证明其有效性。在此基础上,进一步设计了QV-Ka优化方案,并经实验验证。本工作对QKV机制的可解释性分析,为大语言模型架构的未来发展奠定了坚实基础。

原文摘要 · Abstract (English)

Starting from first principles and a linguistic perspective centered on part-of-speech (POS) and syntactic analysis, this paper explores and derives the underlying essence of the Query-Key-Value (QKV) mechanism within the Transformer architecture. Based on this theoretical foundation, we provide a unified explanatory framework for the efficacy of contemporary architectures, including MQA, GQA, and MLA, while identifying their inherent trade-offs and potential optimization trajectories. We introduce the QV paradigm and provide empirical evidence for its validity. Building upon this, we propose the QV-Ka optimization scheme, which is further substantiated through experimental validation. The interpretable theoretical analysis of the QKV mechanism presented in this work establishes a robust foundation for the future evolution of large language model architectures.

注意力机制语言学模型优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。