arXiv:2606.01850cs.AI2026-06被引 1

压缩会破坏大模型的不确定性估计能力,需用新方法评估。

Does Compression Preserve Uncertainty? A Unified Benchmark for Quantized and Sparse LLMs via Conformal Prediction

论文配图:Does Compression Preserve Uncertainty? A Unified Benchmark for Quantized and Sparse LLMs via Conformal Prediction
图 1 · 摘自论文原文
  • 用置信预测统一评测压缩后模型的不确定性
  • 压缩常导致准确率与不确定性脱钩,大模型更抗压
  • 不确定性突增而非渐变,适合安全关键场景评估

量化和剪枝等模型压缩技术广泛用于降低大语言模型(LLM)的部署成本,现有评估几乎仅关注准确性。但在安全关键应用中,模型自我量化不确定性的能力同样重要。我们提出:压缩是否保留了这种能力?通过在五个NLP任务上对12个压缩配置下的LLM进行基准测试,使用置信预测提供无分布假设的不确定性度量。实验发现:(I) 压缩常使准确率与不确定性解耦;(II) 更大的模型比小模型更能吸收压缩带来的不确定性;(III) 不确定性膨胀常呈阈值状而非渐进。结果表明,仅以准确率为标准不足以评估压缩模型的部署准备度,不确定性感知的基准评测应成为压缩流程的标准组成部分。

原文摘要 · Abstract (English)

Model compression techniques such as quantization and pruning are widely used to reduce the deployment cost of large language models (LLMs), with existing evaluations focusing almost exclusively on accuracy preservation. However, in safety-critical applications, a model's ability to reliably quantify its own uncertainty is equally important. We ask: does compression preserve this ability? To answer this question, we benchmark 12 LLMs under various compression configurations across five NLP tasks, using conformal prediction to provide a rigorous, distribution-free measure of uncertainty. Our experiments reveal that: (I) compression frequently decouples accuracy from uncertainty; (II) larger models absorb compression-induced uncertainty far more effectively than smaller ones; and (III) uncertainty inflation is often threshold-like rather than gradual. These results suggest that accuracy-only evaluation is insufficient for assessing the deployment readiness of compressed LLMs, and that uncertainty-aware benchmarking should be a standard component of model compression pipelines.

模型压缩不确定性置信预测大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。