轻量旋转方案与硬件协同,实现4比特大模型高效精准推理。
LightRot: A Light-Weighted Rotation Scheme and Architecture for Accurate Low-Bit Large Language Model Inference

- 采用分组局部旋转与异常方向对齐算法降低量化误差。
- 28nm芯片实现4比特推理峰值能效27.4 TOPS/W,领先现有设计。
- 专为LLaMA2/3等先进模型优化,适用于真实对话场景。
随着大语言模型在多个领域展现出卓越能力,实现节能且高精度的推理变得愈发关键。本文提出LightRot,一种轻量级旋转方案与专用硬件加速器,用于低比特大语言模型推理。该架构融合分组局部旋转(GLR)与异常方向对齐(ODA)算法,并采用基于快速哈达玛变换(FHT)的分层旋转单元,有效缓解低比特量化中的旋转操作能耗问题。所设计加速器在28nm CMOS工艺下实现4比特推理峰值能效27.4 TOPS/W,超越现有最先进方案。不同于依赖高精度推理或仅在基础语言建模任务(如GPT-2)上评估的传统方法,LightRot专为LLaMA2-13B和LLaMA3-8B等先进模型优化,其性能在MT-Bench上得到验证,展现出对真实对话场景的强大适用性,重新定义了基于聊天的AI系统基准。通过算法创新与硬件效率的协同,本工作为可扩展的低比特大模型推理树立新范式,推动可持续人工智能发展。
原文摘要 · Abstract (English)
As large language models (LLMs) continue to demonstrate exceptional capabilities across various domains, the challenge of achieving energy-efficient and accurate inference becomes increasingly critical. This work presents LightRot, a lightweight rotation scheme and dedicated hardware accelerator designed for low-bit LLM inference. The proposed architecture integrates Grouped Local Rotation (GLR) and Outlier Direction Aligning (ODA) algorithms with a hierarchical Fast Hadamard Transform (FHT)-based rotation unit to address key challenges in low-bit quantization, including the energy overhead of rotation operations. The proposed accelerator, implemented in a 28nm CMOS process, achieves a peak energy efficiency of 27.4 TOPS/W for 4-bit inference, surpassing prior state-of-the-art designs. Unlike conventional approaches that rely on higher-precision inference or evaluate on basic language modeling tasks like GPT-2, LightRot is optimized for advanced models such as LLaMA2-13B and LLaMA3-8B. Its performance is further validated on MT-Bench, demonstrating robust applicability to real-world conversational scenarios and redefining benchmarks for chat-based AI systems. By synergizing algorithmic innovations and hardware efficiency, this work sets a new paradigm for scalable, low-bit LLM inference, paving the way for sustainable AI advancements.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。