arXiv:2505.21670cs.CLcs.AI2025-05被引 5

剖析大模型异常值成因,提出高效消除方法

Rethinking the Outlier Distribution in Large Language Models: An In-depth Study

  • 分析大规模激活与通道级异常值的形成机制
  • 提出方法可大幅减少异常值且基本不损失精度
  • 适合关注模型压缩与边缘部署的研究者

研究大语言模型(LLMs)中的异常值至关重要,因其显著影响量化与压缩等性能。异常值常导致严重量化误差,降低模型表现。识别并处理这些异常值可提升量化精度与效率,便于在边缘设备或专用硬件上部署。已有研究识别出两类常见异常值:大规模激活和通道级异常值。尽管已有多种量化算法缓解其影响,但对异常值根源的深入探究仍不足。本文系统研究了这两类异常值的生成机制,并提出相应缓解策略。最终,我们提出若干高效方法,在几乎不影响精度的前提下,有效消除绝大多数大规模激活和通道级异常值。

原文摘要 · Abstract (English)

Investigating outliers in large language models (LLMs) is crucial due to their significant impact on various aspects of LLM performance, including quantization and compression. Outliers often cause considerable quantization errors, leading to degraded model performance. Identifying and addressing these outliers can enhance the accuracy and efficiency of the quantization process, enabling smoother deployment on edge devices or specialized hardware. Recent studies have identified two common types of outliers in LLMs: massive activations and channel-wise outliers. While numerous quantization algorithms have been proposed to mitigate their effects and maintain satisfactory accuracy, few have thoroughly explored the root causes of these outliers in depth. In this paper, we conduct a comprehensive investigation into the formation mechanisms of these outliers and propose potential strategies to mitigate their occurrence. Ultimately, we introduce some efficient approaches to eliminate most massive activations and channel-wise outliers with minimal impact on accuracy.

大模型量化异常值

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。