针对量化后保持模型不确定性,提出动态校准数据选择方法
Target-Aware Calibration Data Selection for Preserving Uncertainty in Quantized Language Models

- 基于目标定制化选择高疑虑样本与通用锚点构建校准集
- 在8个模型、9个任务上验证,不同目标需不同校准策略
- 适合关注推理可靠性与不确定性的模型部署场景
量化广泛用于大语言模型部署,但其对置信度、边界和拒答等不确定性行为的影响常被忽视。本文将量化校准数据选择问题建模为依赖目标的不确定性保留问题。不同应用场景关注输入分布的不同区域,而以往方法多聚焦准确率优化或量化后调校分数。我们通过分布与边界保留风险形式化该目标,并提出混合不匹配论据:单一校准方案无法适配所有目标。引入轻量级预量化方法DPQ,利用全精度预测构造与目标对齐的高疑虑样本与通用锚点混合集。在8个语言模型、9个NLP基准和22种对比方法中,最优固定配方随目标变化:DPQ-r75在SQuAD2答案可回答性边界保留上表现最佳,而更温和或单信号变体(如DPQ-r50、仅置信度、仅熵)则更好保留多选问答的整体行为。结果表明,校准数据应根据具体部署需求选择,而非作为固定量化细节处理。
原文摘要 · Abstract (English)
Quantization is widely used to deploy large language models, but its effect on uncertainty behavior, such as confidence, margins, and abstention, is rarely treated as a primary objective. We frame calibration-data selection for quantization as a target-dependent uncertainty-preservation problem. Different deployments emphasize different regions of the input distribution, yet prior work mainly optimizes accuracy-oriented compression metrics or adjusts scores after quantization. We formalize this goal with distributional and boundary preservation risks, and provide a simple mixture-mismatch argument explaining why no single calibration recipe should be expected to fit all targets. We introduce Doubt-Preserving Quantization (DPQ), a lightweight pre-quantization recipe family that uses full-precision predictions to construct target-aligned calibration mixtures of high-doubt examples and generic anchors. Across 8 language models, 9 NLP benchmarks, and 22 comparison methods, the leading fixed recipe changes with the preservation target: DPQ-r75 leads on SQuAD2 answerability-boundary preservation, while milder or single-signal variants, including DPQ-r50, confidence-only, and entropy-only, better preserve broad multiple-choice QA behavior. These results show that calibration data should be selected for the specific full-precision score behavior a deployment needs to preserve, rather than treated as a fixed quantization detail.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。