量化让大模型偏见翻转,却骗过平均评分。
Investigating Social Bias Changes in Quantized Language Models
- 首次大规模测试50个量化模型的偏见变化。
- 4位量化模型行为改变是8位的4-6倍,偏见可能恶化18.6%。
- 偏见翻转不均等,需针对性评估与干预。
后训练量化虽降低大语言模型内存占用,却以聚合指标无法捕捉的方式改变其社会偏见。我们首次在包含13个封闭与开放题偏见数据集的PostTrainingBiasBench基准上,对50个量化模型进行大规模评估。发现一种称为量化诱导偏见翻转的现象:量化使模型响应在有偏与无偏间来回切换,高达21%的响应发生翻转,但整体偏见得分未变。该现象与模型不确定性强相关,高不确定性的响应翻转概率是自信响应的3-11倍。量化强度加剧此效应,4位量化模型的行为改变量是8位模型的4-6倍。关键的是,这种变化对不同群体影响不对称:某些群体偏见恶化最高达18.6%,另一些则改善最多14.1%,导致整体结果看似中性。大模型并未表现出一致鲁棒性,群体偏见变化在不同模型家族间不可预测。研究揭示压缩本质改变了偏见模式,亟需量化后的系统性评估与干预以保障实际应用可靠性。
原文摘要 · Abstract (English)
Post-training quantization reduces the memory needed to run large language models but alters their social biases in ways that aggregate metrics fail to capture. We present the first large-scale study of 50 quantized models evaluated on PostTrainingBiasBench, a unified benchmark of 13 closed- and open-ended bias datasets. We identify a phenomenon we term quantization-induced bias flipping, in which quantization causes models to change responses from biased to unbiased and vice versa, up to 21% of the time, despite no change in aggregate bias scores. These flips are strongly associated with model uncertainty, where the responses with high uncertainty are 3-11x more likely to change than the confident ones. Quantization strength amplifies this effect, with 4-bit quantized models exhibiting 4-6x more behavioral changes than 8-bit quantized models. Critically, these changes create asymmetric impacts across demographic groups, where bias can worsen by up to 18.6% for some groups while improving by 14.1% for others, yielding misleadingly neutral aggregate outcomes. Larger models show no consistent robustness advantage, and group-specific shifts vary unpredictably across model families. Our findings demonstrate that compression fundamentally alters bias patterns, requiring crucial post-quantization evaluation and interventions to ensure reliability in practice.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。