arXiv:2604.08561cs.CLcs.LG2026-04中稿 · KDD

分析大模型去偏后嵌入空间变化,验证公平性改进是否真实发生。

A Representation-Level Assessment of Bias Mitigation in Foundation Models

论文配图:A Representation-Level Assessment of Bias Mitigation in Foundation Models
图 1 · 摘自论文原文
  • 通过嵌入空间分析,对比基础与去偏模型的性别-职业关联变化。
  • 去偏后性别与职业的嵌入差异显著降低,内部表征更中立均衡。
  • 提出新数据集 WinoDec,支持对解码器类模型的公平性评估。

我们研究了去偏技术如何重塑仅编码器和仅解码器型大模型的嵌入空间,通过表征分析实现对模型行为的内部审计。以 BERT 和 Llama2 为代表架构,比较基准模型与去偏版本在性别与职业词间关联的转变。结果表明,去偏能有效降低嵌入空间中性别-职业的差异,使内部表征趋向中性与平衡。这种表征变化在两类模型中均一致出现,说明公平性提升可表现为可解释的几何变换。该发现将嵌入分析确立为验证去偏方法有效性的重要工具。为进一步推动对解码器模型的评估,我们构建了包含 4,000 条序列的 WinoDec 数据集,并公开发布(https://github.com/winodec/wino-dec)。

原文摘要 · Abstract (English)

We investigate how successful bias mitigation reshapes the embedding space of encoder-only and decoder-only foundation models, offering an internal audit of model behaviour through representational analysis. Using BERT and Llama2 as representative architectures, we assess the shifts in associations between gender and occupation terms by comparing baseline and bias-mitigated variants of the models. Our findings show that bias mitigation reduces gender-occupation disparities in the embedding space, leading to more neutral and balanced internal representations. These representational shifts are consistent across both model types, suggesting that fairness improvements can manifest as interpretable and geometric transformations. These results position embedding analysis as a valuable tool for understanding and validating the effectiveness of debiasing methods in foundation models. To further promote the assessment of decoder-only models, we introduce WinoDec, a dataset consisting of 4,000 sequences with gender and occupation terms, and release it to the general public. (https://github.com/winodec/wino-dec)

大模型去偏嵌入分析公平性评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。