arXiv:2606.17506cs.CL2026-06

用哲学推理任务检测大模型判断偏见时的隐性偏见。

Evaluating Second-Order Bias of LLMs Through Epistemic Entitlement

  • 基于认知正当性理论设计逻辑推理任务,评估模型对偏见内容的评判标准。
  • 发现模型在无充分依据时仍会根据身份标签推断内容可接受性,且差异随目标群体变化。
  • 适用于关注模型判别偏见能力、需深入理解隐性偏见的研究者与开发者。

当前大模型偏见评估主要关注其是否生成或暗示偏见内容,但当大模型被用作偏见裁判时,可能以更隐蔽的方式表现出社会偏见,即第二阶偏见——对偏见内容判断中的社会偏见。本文提出一种基于正当性认识论的哲学驱动推理任务,将偏见视为误导性基础知识,并设计任务让模型判断某偏见文本对特定群体是否可接受。我们开发两个简单度量指标:衡量模型在缺乏足够支持时对可接受性的身份推断程度,以及该推断在不同目标群体间的差异。在开放与闭源模型上的评估表明,该任务能绕过安全防护机制,揭示出系统性偏见,反映隐含社会认知图景,且模型仍受身份标签触发。研究强调需在判别任务中评估大模型偏见,并倡导采用更理论化的偏见评估方法。代码与模型响应已公开于 https://github.com/uofthcdslab/second-order-bias。

原文摘要 · Abstract (English)

Evaluations of social bias in LLMs largely focus on whether models generate or imply biased content. However, as LLMs are increasingly used as judges of bias, they may exhibit social biases in subtler ways in how they evaluate biased content, which current methods do not systematically capture. We call this second-order bias: social bias in an LLM's judgment about social bias, which we evaluate through a novel, philosophically grounded reasoning task. Drawing on entitlement epistemology, we conceptualize bias as misplaced foundational knowledge that shapes an agent's rational inquiry, and derive a logical reasoning task for LLMs to judge to whom a biased text is acceptable or non-acceptable. We develop two simple metrics to measure how biased LLM judges are in inferring demographics for acceptability without sufficient support, and how these inferences vary across groups targeted by biased texts. Evaluating open and closed models, we find that our task evades safety guardrails by surfacing bias in model judgment. It varies systematically across target groups, reflects implicit social maps, and shows how models are still triggered by demographic labels. Our work points to the need for LLM bias evaluation in judgment tasks and broadly, for more theoretically grounded approaches to bias evaluation in NLP. We release our code and model responses at https://github.com/uofthcdslab/second-order-bias.

大模型偏见推理评估认知机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。