arXiv:2606.07964cs.CL2026-06

PCA去偏方法只移除部分性别偏见,还破坏语义结构。

What Does Debiasing Really Remove? A Geometric Study of PCA-Based Gender Debiasing in Word Embeddings

论文配图:What Does Debiasing Really Remove? A Geometric Study of PCA-Based Gender Debiasing in Word Embeddings
图 1 · 摘自论文原文
  • 用主成分分析定位并移除嵌入空间中的性别偏见方向。
  • 移除越多主成分,语义关系越受损,导致几何失真。
  • 去偏效果因指标不同而异,无通用最优方案。

基于主成分分析(PCA)的去偏方法广泛用于降低大语言模型中词向量的性别偏见,但其实际移除的内容及破坏性仍不明确。本文系统研究了PCA-based性别去偏的几何特性。实验表明,直接性别偏见主要集中在第一主成分,支持低秩偏见假设;然而,通过WEAT衡量的关联偏见并不集中于主成分方向,而是分散在多个维度。此外,随着移除主成分数量增加,嵌入空间的几何结构持续退化,影响语义关系和向量间语义结构。结果表明,去偏过程存在权衡:虽能有效减少部分直接偏见,却无法消除分布式的关联偏见,并引入几何畸变。且不存在统一最优去偏程度,平衡取决于具体度量方式与嵌入类型。整体揭示词向量偏见非纯低秩,简单子空间移除不足以实现全面去偏。

原文摘要 · Abstract (English)

Debiasing methods based on principal component analysis (PCA) are broadly used to reduce gender bias in word embeddings used in LLMs, yet it remains unclear what aspects of bias they actually remove and how destructive this process is. These methods are based on the understanding that bias resides in a low-dimensional subspace, with the assumption that most of it can be captured by a few principal components. In this work, we conduct a systematic geometric analysis of PCA-based gender debiasing and investigate what is actually removed from the embedding space. Our experiments across multiple embeddings show that direct gender bias is primarily concentrated in the first principal component, supporting the low-rank bias hypothesis. However, associative bias measured by WEAT does not align with these principal directions and is instead spread across multiple embedding dimensions. Furthermore, as expected, we demonstrate that removing an increasing number of principal components leads to a consistent degradation of the embedding geometry, affecting semantic structure and vector relationships. These results reveal that PCA-based debiasing operates as a trade-off: while it effectively reduces certain forms of direct bias, it fails to eliminate distributed associations and introduces geometric distortion. Moreover, there is no universal optimal level of debiasing, as the balance between bias reduction and semantic preservation depends on the chosen metric and embedding. Overall, our findings suggest that bias in word embeddings is not purely low-rank and that simple subspace removal methods may be insufficient for comprehensive debiasing.

去偏词向量主成分分析语义结构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。