arXiv:2505.15524cs.CLcs.AI2025-05被引 5

无需人工标注,通过向量空间结构检测大模型隐含偏见。

Evaluate Bias without Manual Test Sets: A Concept Representation Perspective for LLMs

  • 利用概念激活向量与稀疏自编码器提取可解释概念表征。
  • 在无标签数据下与传统评估方法相关性超0.85,效果一致。
  • 可发现难检测的隐性偏见,适合模型公平性研究者使用。

大语言模型中的偏见严重威胁其可靠性与公平性。本文关注一种常见偏见:当模型中两个参考概念(如情感极性“正”与“负”)与第三个目标概念(如评价维度)存在不对称关联时,模型产生非预期偏见。例如,“食物”的理解不应偏向某种情感。现有评估方法依赖人工构建社会群体标签数据集,成本高且覆盖范围有限。为此,本文提出无测试集偏见分析框架BiasLens,结合概念激活向量(CAVs)与稀疏自编码器(SAEs),提取可解释的概念表征,并通过测量目标概念与各参考概念间表征相似性的差异来量化偏见。即使无标注数据,BiasLens与传统评估指标具有强一致性(斯皮尔曼相关系数r > 0.85)。此外,该方法揭示了传统方法难以检测的偏见形式,例如在模拟临床场景中,患者保险状态会引发诊断评估偏见。总体而言,BiasLens为偏见发现提供了一种可扩展、可解释、高效的范式,推动大模型公平性与透明度提升。

原文摘要 · Abstract (English)

Bias in Large Language Models (LLMs) significantly undermines their reliability and fairness. We focus on a common form of bias: when two reference concepts in the model's concept space, such as sentiment polarities (e.g., "positive" and "negative"), are asymmetrically correlated with a third, target concept, such as a reviewing aspect, the model exhibits unintended bias. For instance, the understanding of "food" should not skew toward any particular sentiment. Existing bias evaluation methods assess behavioral differences of LLMs by constructing labeled data for different social groups and measuring model responses across them, a process that requires substantial human effort and captures only a limited set of social concepts. To overcome these limitations, we propose BiasLens, a test-set-free bias analysis framework based on the structure of the model's vector space. BiasLens combines Concept Activation Vectors (CAVs) with Sparse Autoencoders (SAEs) to extract interpretable concept representations, and quantifies bias by measuring the variation in representational similarity between the target concept and each of the reference concepts. Even without labeled data, BiasLens shows strong agreement with traditional bias evaluation metrics (Spearman correlation r > 0.85). Moreover, BiasLens reveals forms of bias that are difficult to detect using existing methods. For example, in simulated clinical scenarios, a patient's insurance status can cause the LLM to produce biased diagnostic assessments. Overall, BiasLens offers a scalable, interpretable, and efficient paradigm for bias discovery, paving the way for improving fairness and transparency in LLMs.

偏见检测可解释性大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。