用光电子器件替代传统Softmax,实现超低延迟的注意力计算。
Integrated electro-optic attention nonlinearities for transformers
- 用钽酸锂调制器实现模拟非线性运算,替代数字Softmax和Sigmoid。
- 在4比特量化下仍保持高精度,推理延迟显著降低。
- 适合需要高速低功耗计算的Transformer模型部署场景。
Transformer已成为语言处理与计算机视觉的主流架构,其核心是依赖Softmax函数的非线性、非负映射注意力机制。尽管Softmax操作仅占总运算量不足1%,却可能严重拖慢整体推理延迟。本文采用薄膜铌酸锂(TFLN)马赫-曾德尔调制器(MZM)作为模拟非线性计算单元,大幅降低非线性计算延迟。我们实现了电光版Softmax与Sigmoid的替代,并在视觉Transformer与大语言模型中评估性能。系统在极端4比特输入输出量化下仍保持高准确率。进一步在最高10 GBaud编码速率下表征系统噪声,并评估不同噪声条件下的模型鲁棒性。结果表明,TFLN调制器可作为混合共封装硬件中的非线性函数单元,实现高速、低功耗的非线性计算。
原文摘要 · Abstract (English)
Transformers have emerged as the dominant neural-network architecture, achieving state-of-the-art performance in language processing and computer vision. At the core of these models lies the attention mechanism, which requires a nonlinear, non-negative mapping using the Softmax function. However, although Softmax operations account for less than 1% of the total operation count, they can disproportionately bottleneck overall inference latency. Here, we use thin-film lithium niobate (TFLN) Mach-Zehnder modulators (MZMs) as analog nonlinear computational elements to drastically reduce the latency of nonlinear computations. We implement electro-optic alternatives to digital Softmax and Sigmoid, and evaluate their performance in Vision Transformers and Large Language Models. Our system maintains highly competitive accuracy, even under aggressive 4-bit input-output quantization of the analog units. We further characterize system noise at encoding speeds up to 10 GBaud and assess model robustness under various noise conditions. Our findings suggest that TFLN modulators can serve as nonlinear function units within hybrid co-packaged hardware, enabling high-speed and energy-efficient nonlinear computation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。