提出注意力采样器,高效处理大规模注意力模型的流式数据
Towards Sampling Data Structures for Tensor Products in Turnstile Streams
- 基于重要性采样设计新型注意力采样器
- 理论证明空间与更新时间复杂度显著降低
- 适用于多种模型架构,适合实时大数据场景
本文研究人工智能中大规模注意力模型在流式数据下的计算挑战,结合经典ℓ₂采样定义与大语言模型注意力机制的最新进展,提出注意力采样器的新定义。该方法显著降低了传统注意力机制的计算负担。我们从理论上分析了注意力采样器的有效性,涵盖空间占用和更新时间。此外,该框架展现出良好的可扩展性与跨模型架构、多领域的广泛适用性。
原文摘要 · Abstract (English)
This paper studies the computational challenges of large-scale attention-based models in artificial intelligence by utilizing importance sampling methods in the streaming setting. Inspired by the classical definition of the $\ell_2$ sampler and the recent progress of the attention scheme in Large Language Models (LLMs), we propose the definition of the attention sampler. Our approach significantly reduces the computational burden of traditional attention mechanisms. We analyze the effectiveness of the attention sampler from a theoretical perspective, including space and update time. Additionally, our framework exhibits scalability and broad applicability across various model architectures and domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。