arXiv:2511.06160cs.AIcs.CL2025-11ACL被引 1

用逻辑谜题检测大模型推理中的隐性偏见,发现符合性别刻板印象的答案更易被采纳。

Evaluating Implicit Biases in LLM Reasoning through Logic Grid Puzzles

  • 设计逻辑谜题框架PRIME,通过三类变体对比刻板、反刻板与中性情境下的推理表现。
  • 实验显示模型在符合性别刻板印象的解法上准确率显著更高,揭示隐性偏见存在。
  • 适合关注大模型公平性、推理偏差评估的研究者与安全开发者参考。

尽管近期的安全防护机制能抑制明显的偏见输出,但在复杂逻辑推理任务中仍会浮现更隐蔽的社会偏见,而现有评估基准难以捕捉此类问题。为此,我们提出新的评估框架PRIME(Puzzle Reasoning for Implicit Biases in Model Evaluation),利用逻辑网格谜题系统性地探测社会刻板印象对大模型推理与决策的影响。该方法支持自动生成与验证,并可调节复杂度与偏见设置。PRIME包含从同一谜题结构衍生出的刻板、反刻板和中性版本,实现可控且精细的对比分析。我们在多个模型家族上测试不同规模谜题,并评估基于提示的缓解策略效果。聚焦于性别刻板印象,结果表明模型在解法符合刻板印象时推理准确率更高,凸显了PRIME在诊断与量化大模型演绎推理中社会偏见方面的价值,尤其在强调公平性的场景中至关重要。

原文摘要 · Abstract (English)

While recent safety guardrails effectively suppress overtly biased outputs, subtler forms of social bias emerge during complex logical reasoning tasks that evade current evaluation benchmarks. To fill this gap, we introduce a new evaluation framework, PRIME (Puzzle Reasoning for Implicit Biases in Model Evaluation), that uses logic grid puzzles to systematically probe the influence of social stereotypes on logical reasoning and decision making in LLMs. Our use of logic puzzles enables automatic generation and verification, as well as variability in complexity and biased settings. PRIME includes stereotypical, anti-stereotypical, and neutral puzzle variants generated from a shared puzzle structure, allowing for controlled and fine-grained comparisons. We evaluate multiple model families across puzzle sizes and test the effectiveness of prompt-based mitigation strategies. Focusing our experiments on gender stereotypes, our findings highlight that models consistently reason more accurately when solutions align with stereotypical associations. This demonstrates the significance of PRIME for diagnosing and quantifying social biases perpetuated in the deductive reasoning of LLMs, where fairness is critical.

模型偏见逻辑推理评估框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。