用ZCA白化校准词向量空间,让偏见检测更可靠
Calibrating WEAT Against Anisotropy: ZCA Whitening as a Geometric Pre-Processing Step for Embedding Association Tests

- 用ZCA白化修复词向量空间的各向异性问题
- 70个模型-任务组合中,30%以上结果显著性改变
- 适合关注模型偏见评估可信度的研究者
本文提出使用零相位成分分析(ZCA)白化作为词向量关联测试(WEAT)的几何预处理步骤。WEAT广泛用于计算社会科学和人工智能公平性研究中的偏见测量,依赖余弦相似度作为语义关联指标,该方法假设词向量空间近似各向同性。然而,先前研究表明许多主流语言模型不满足此假设,引发对偏见测量可靠性担忧。ZCA白化在最小扰动原向量的前提下,将嵌入空间协方差转化为单位矩阵,恢复WEAT所依赖的各向同性条件。我们在10个标准WEAT测试集、7种跨越三种架构的模型上进行评估,共涵盖70个模型-任务组合。结果表明,ZCA白化显著降低所有模型的各向异性;对于高度各向异性的模型,语义相似性基准性能进一步提升,说明校准后的空间更准确捕捉语义关联。校准后,超过30%的WEAT结果显著性状态发生改变,效应量在不同偏见类别中双向变化。这表明未校准测量可能同时高估和低估嵌入空间中的关联。因此,此前在各向异性空间中报告的偏见测量应谨慎解读,并建议用校准方法重新评估。本方法为恢复WEAT在计算社会科学与AI公平性研究中的测量基础作出贡献。
原文摘要 · Abstract (English)
We propose Zero-phase Component Analysis (ZCA) whitening as a geometric pre-processing step for the Word Embedding Association Test (WEAT). WEAT is a bias measurement method widely used in both computational social science and AI fairness research. It relies on cosine similarity as a measure of semantic association, which assumes that the embedding space is approximately isotropic. However, prior work has reported that many widely used language models do not satisfy this assumption, raising concerns about the reliability of bias measurements. ZCA whitening transforms the covariance of the embedding space into the identity matrix while minimizing perturbation to the original vectors. This transformation restores the isotropy condition on which WEAT relies. We evaluate our approach on ten standard WEAT test suites and seven models spanning three architectural families, yielding 70 model-task combinations. The results show that ZCA whitening substantially reduces the anisotropy of the embedding spaces across all models. Particularly for highly anisotropic models, we further observe improvements on standard semantic similarity benchmarks, indicating that the calibrated space better captures semantic associations. After calibration, over 30% of WEAT results change significance status, and effect sizes shift in both directions depending on bias category. These shifts suggest that uncalibrated measurements may both overestimate and underestimate the associations encoded in the embedding space. These findings indicate that previously reported bias measurements in anisotropic embedding spaces should be interpreted with caution and may benefit from re-evaluation with calibrated methods. Our approach contributes to restoring the measurement foundation of WEAT across both computational social science and AI fairness research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。