arXiv:2505.13963cs.CLcs.LG2025-05被引 3

量化会损失事实记忆能力,但部分方法仍能保持效果。

Through a Compressed Lens: Investigating The Impact of Quantization on Factual Knowledge Recall

  • 测试三种量化方法在不同比特位下的表现
  • 小模型量化后事实回忆能力下降更明显
  • BitSandBytes在保持知识召回上表现最佳

量化技术被广泛用于加速大语言模型(LLM)推理与部署。尽管量化对各类模型能力的影响已有广泛研究,但其对事实知识召回(FKR)——即模型调用存储知识的能力——的影响仍缺乏深入探讨。为此,我们使用三种常见量化技术,在不同比特位下进行系统实验,并结合可解释性分析,考察知识记忆与潜在多跳推理两个任务。结果表明,量化通常导致模型内部信息丢失,从而削弱其事实知识召回能力,尤其在同架构中较小的模型表现更差。然而,低比特量化并非总是导致性能下降,某些情况下甚至可能提升知识召回。其中,BitSandBytes在保持原全精度模型的FKR方面表现最优。尽管存在模型与方法间的差异,量化整体仅造成轻微性能退化,仍是一种有效的压缩策略。

原文摘要 · Abstract (English)

Quantization methods are widely used to accelerate inference and streamline the deployment of large language models (LLMs). Although quantization's effects on various LLM capabilities have been extensively studied, one critical area remains underexplored: factual knowledge recall (FKR), the process by which LLMs access stored knowledge. To this end, we conduct comprehensive experiments using three common quantization techniques at distinct bit widths, in conjunction with interpretability-driven analyses on two tasks, knowledge memorization and latent multi-hop reasoning. We show that quantization typically results in information loss within LLMs, consequently diminishing their capacity for FKR. This effect is particularly amplified in smaller models within the same architectural families. However, models quantized at reduced bit precision do not consistently exhibit inferior performance and occasionally quantization may even enhance model FKR. We find that BitSandBytes demonstrates highest preservation of the original full-precision model's FKR. Despite variability across models and methods, quantization causes modest performance degradation and remains an effective compression strategy.

量化知识召回大模型压缩

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。