arXiv:2601.00282cs.CLcs.AI2026-01被引 2

量化会降低大模型自解释质量,需针对性验证。

Can Large Language Models Still Explain Themselves? Investigating the Impact of Quantization on Self-Explanations

  • 测试三种量化方法在不同比特下的自解释能力
  • 自解释质量下降最高达4.4%,可信度下降8.5%
  • 大模型对量化更敏感,且无万能量化方案

量化广泛用于加速大语言模型(LLM)推理和部署,但其对自解释(SEs)的影响尚未明确。自解释是模型生成的用于说明自身输出的理由,涉及对决策过程的推理,可能对量化特别敏感。随着自解释在高风险场景中越来越重要,理解量化是否以及在多大程度上损害自解释的质量与忠实性至关重要。为此,我们研究了两种自解释类型:自然语言解释(NLEs)和反事实例子,使用三种常见量化技术在不同位宽下进行测试。结果表明,量化通常导致自解释质量下降最多4.4%,忠实性下降最多3.9%。用户研究进一步显示,量化显著降低自解释的连贯性和可信度,降幅高达8.5%。相比小模型,大模型在自解释质量上对量化更不耐受,但在忠实性上保持相对稳定。此外,没有一种量化技术在任务准确率、自解释质量和忠实性上始终表现最佳。由于量化影响因情境而异,且在特定情况下可能显著,我们建议针对具体应用场景验证自解释质量。尽管存在明显下降,只要经过恰当验证,量化仍是有效的压缩手段。

原文摘要 · Abstract (English)

Quantization is widely used to accelerate inference and streamline the deployment of large language models (LLMs), yet its effects on self-explanations (SEs) remain unexplored. SEs, generated by LLMs to justify their own outputs, require reasoning about the model's own decision-making process, a capability that may exhibit particular sensitivity to quantization. As SEs are increasingly relied upon for transparency in high-stakes applications, understanding whether and to what extent quantization degrades SE quality and faithfulness is critical. To address this gap, we examine two types of SEs: natural language explanations (NLEs) and counterfactual examples, generated by LLMs quantized using three common techniques at distinct bit widths. Our findings indicate that quantization typically leads to moderate declines in both SE quality (up to 4.4%) and faithfulness (up to 3.9%). The user study further demonstrates that quantization considerably diminishes both the coherence and trustworthiness of SEs (by up to 8.5%). Compared to smaller models, larger models show limited resilience to quantization in terms of SE quality but maintain more faithfulness. Moreover, no quantization technique consistently excels across task accuracy, SE quality, and faithfulness. Because quantization's impact varies considerably by context and can be sizable in specific cases, we recommend validating SE quality for the intended use case. Despite these sometimes considerable drops, quantization remains an effective compression technique when its impact on SEs is properly validated.

大模型量化自解释可信度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。