arXiv:2604.13806cs.LG2026-04

用稳定曲率估计提升低比特量化精度,小样本下表现更稳健。

Robust Ultra Low-Bit Post-Training Quantization via Stable Diagonal Curvature Estimate

论文配图:Robust Ultra Low-Bit Post-Training Quantization via Stable Diagonal Curvature Estimate
图 1 · 摘自论文原文
  • 基于对角海森矩阵与加权最小二乘法,过滤噪声干扰。
  • 在超低比特下平均提升零样本准确率7.01%,最高达14.01%。
  • 适合资源受限场景,仅需少量校准数据即可稳定部署。

大语言模型(LLMs)广泛应用,但规模大导致部署困难。后训练量化(PTQ)通过小规模校准集实现无重训练的内存压缩。现有基于海森矩阵的PTQ方法依赖跨通道依赖来补偿量化误差,但在低比特时因校准数据少导致曲率估计噪声大而性能下降。我们提出DASH-Q框架,采用对角海森近似和迭代加权最小二乘法。通过舍弃易受噪声影响的依赖关系,DASH-Q在保留关键特征能量的同时滤除采样噪声。在五个基准大模型上,相比最强基线,平均提升零样本准确率7.01%,最高提升14.01%,且在极小校准数据下仍保持鲁棒稳定性能。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are widely used across many domains, but their scale makes deployment challenging. Post-Training Quantization (PTQ) reduces memory footprint without retraining by leveraging a small calibration set. Recent Hessian-based PTQ methods compensate quantization error via cross-channel dependencies, but such approaches degrade at low bit-widths due to noisy curvature estimates from limited calibration data. We propose DASH-Q, a robust PTQ framework using diagonal Hessian approximation and iterative weighted least squares. By discarding noise-prone dependencies, DASH-Q filters sampling noise while prioritizing the preservation of salient feature power. We outperform other PTQ baselines in ultra low-bit regime, improving zero-shot accuracy by 7.01% on average and up to 14.01% over the strongest baselines across five baseline LLM models, while showing robust and stable performance with very small calibration data.

量化大模型低比特

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。