arXiv:2601.12033cs.CL2026-01ACL

量化会损害大模型公平性与安全性,该研究提出保护关键权重来解决。

Preserving Fairness and Safety in Quantized LLMs Through Critical Weight Protection

  • 识别并保护影响公平与安全的关键权重,避免量化时丢失。
  • 动态量化比静态量化更稳定,但非英语语言安全性下降更严重。
  • 无需重训练即可提升模型可信度,适合部署在多语言场景。

量化被广泛用于降低大语言模型的计算开销;然而,其对公平性和安全性的影响,尤其是在动态量化和多语言情境下的影响仍不明确。本文系统研究了静态与动态量化方法在衡量内在与外在偏见及安全对齐的基准上的表现。公平性评估覆盖英语、法语、荷兰语、西班牙语和土耳其语;安全性聚焦英语、韩语和阿拉伯语。结果表明,量化会持续降低公平性与安全性,动态量化方法稳定性优于静态方法。不同语言间公平性退化程度不同,而非英语环境下安全性下降尤为显著。为此,我们提出关键权重保护(Critical Weight Protection)技术,在量化过程中识别并保留对公平性与安全性至关重要的权重。该方法有效缓解偏差与安全问题,无需昂贵的微调或对齐过程,兼顾可信性与效率。

原文摘要 · Abstract (English)

Quantization is widely adopted to reduce the computational cost of large language models (LLMs); however, its implications for fairness and safety, particularly in dynamic quantization and multilingual contexts, remain underexplored. In this work, we conduct a systematic study of how static and dynamic quantization methods impact fairness and safety across benchmarks measuring intrinsic and extrinsic bias and safety alignment. For fairness, we evaluate English, French, Dutch, Spanish, and Turkish; for safety, we focus on English, Korean, and Arabic. Our findings reveal that quantization consistently degrades fairness and safety, with dynamic methods demonstrating greater stability than static ones. Moreover, fairness degradation varies across languages, while safety deterioration is especially pronounced in non-English settings. To address these risks, we introduce Critical Weight Protection, a novel technique that identifies and preserves fairness- and safety-critical weights during quantization. This approach effectively mitigates bias and safety deterioration without costly retraining or alignment, maintaining trustworthiness while retaining efficiency.

量化公平性安全性LLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。