量化不仅能提速,还能让视觉语言模型更可靠。
Less Precise Can Be More Reliable: A Systematic Evaluation of Quantization's Impact on VLMs Beyond Accuracy
- 通过量化抑制高秩特征,强化低秩稳健特征。
- 量化后模型在准确率、校准度、异常检测上均提升。
- 适合追求高效可靠的视觉语言模型部署者。
视觉语言模型(如CLIP)在零样本分类和关键任务(如分布外检测)中表现卓越,但其高计算成本限制了实际部署。尽管量化是提升效率的常用方法,其对准确性之外的可靠性指标的影响仍缺乏系统研究。本研究对多种配置下的视觉语言模型量化进行了大规模评估,覆盖超70万次实验。结果表明,量化虽引入噪声,却能同时提升准确率、校准度、分布外检测能力及抗噪鲁棒性,但对协变量偏移或伪相关无改善。我们发现,量化通过抑制高秩谱成分,迫使模型依赖更稳健的低秩特征,从而实现泛化与抗噪能力增强。这一谱滤波机制揭示了量化超越传统正则化的深层作用,为以量化实现更快、更可靠的视觉语言模型部署提供了新路径。
原文摘要 · Abstract (English)
Vision-Language Models (VLMs) such as CLIP have revolutionized zero-shot classification and safety-critical tasks, including Out-of-Distribution (OOD) detection. However, their high computational cost hinders efficient real-world deployment. While quantization is a standard solution for efficiency, its broader impact on reliability metrics beyond simple Top-1 accuracy remains critically under-explored. In this study, we conduct a large-scale evaluation of VLM quantization across a comprehensive experimental suite of over 700k evaluation runs with varying configurations. We find that, contrary to the assumption that quantization's noise degrades performance, it can simultaneously improve accuracy, calibration, OOD detection, and robustness to noise, though not to covariate shift or spurious correlations. We leverage these counterintuitive findings to characterize the mechanics of quantization beyond simple regularization: we show that quantization dampens high-rank spectral components, compelling the model to rely more heavily on robust, low-rank features. Ultimately, this spectral filtering effect drives the observed improvements in generalization and noise tolerance, establishing a pathway to deploy faster, more reliable VLMs by utilizing quantization beyond its conventional role.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。