arXiv:2410.20265cs.CVcs.CY2024-10被引 3

量化压缩视觉语言模型,但社会偏见变化不一致,无法预测。

You Never Know: Quantization Induces Inconsistent Biases in Vision-Language Foundation Models

  • 测试四种量化设置,评估三类CLIP模型在三个数据集上的表现。
  • 量化未导致偏见放大或缩小的系统性变化,方向和程度均不稳定。
  • 适合关注模型压缩公平性、部署风险的研究者与工程师。

我们研究了压缩视觉语言基础模型的常用技术——量化,对模型生成社会公平输出能力的影响。与单模态模型中压缩会一致放大社会偏见的先前发现不同,我们在三种CLIP变体、三个数据集及四种量化设置下的广泛评估显示:尽管个别模型表现出偏见,但量化并未在一组压缩模型中引发偏见幅度或方向的稳定变化。这一结果挑战了量化必然加剧偏见的假设。

原文摘要 · Abstract (English)

We study the impact of a standard practice in compressing foundation vision-language models - quantization - on the models' ability to produce socially-fair outputs. In contrast to prior findings with unimodal models that compression consistently amplifies social biases, our extensive evaluation of four quantization settings across three datasets and three CLIP variants yields a surprising result: while individual models demonstrate bias, we find no consistent change in bias magnitude or direction across a population of compressed models due to quantization.

量化偏见分析CLIP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。