arXiv:2602.04742cs.CYcs.CL2026-02被引 1

推理能力可显著降低大模型的隐性社会偏见

Inference-Time Reasoning Selectively Reduces Implicit Social Bias in Large Language Models

  • 通过启用推理机制,减少模型在15个刻板印象话题上的隐性偏见
  • 推理使隐性偏见下降,但对非社会类隐性关联无影响
  • 适合关注模型公平性与认知理论融合的研究者

基于心理学中的显性与隐性偏见区分,已有研究发现,尽管大语言模型(LLMs)经过后训练对齐和安全处理以避免显性偏见表达,但在类似内隐联想测试(IAT)的间接任务中仍表现出显著隐性偏见。最近研究表明,推理过程会削弱依赖隐性统计学习的任务表现。受人类认知中隐性联想与统计学习之间理论关联的启发,本文考察了推理型推断对LLM隐性偏见的影响。结果发现,在15个刻板印象主题上,部分模型类别在启用推理后,其隐性偏见测量值显著降低。该效应具有领域特异性:对非社会类隐性关联无类似改善。随着推理功能越来越多地被默认启用,这些发现表明其可能显著改变某些系统的公平性评估结果,并引发关于对齐程序与推理机制如何共同影响偏见缓解的思考。更广泛地,本工作展示了认知科学与心理学理论如何为人工智能评估研究提供方法论与解释框架,揭示模型行为的新洞见。

原文摘要 · Abstract (English)

Drawing on constructs from psychology, prior work has identified a distinction between explicit and implicit bias in large language models (LLMs). While many LLMs undergo post-training alignment and safety procedures to avoid expressions of explicit social bias, they still exhibit significant implicit biases on indirect tasks resembling the Implicit Association Test (IAT). Recent work has further shown that inference-time reasoning can impair LLM performance on tasks that rely on implicit statistical learning. Motivated by a theoretical link between implicit associations and statistical learning in human cognition, we examine how reasoning-enabled inference affects implicit bias in LLMs. We find that enabling reasoning significantly reduces measured implicit bias on an IAT-style evaluation for some model classes across fifteen stereotype topics. This effect appears specific to social bias domains, as we observe no corresponding reduction for non-social implicit associations. As reasoning is increasingly enabled by default in deployed LLMs, these findings suggest that it can meaningfully alter fairness evaluation outcomes in some systems, while also raising questions about how alignment procedures interact with inference-time reasoning to drive variation in bias reduction across model types. More broadly, this work highlights how theory from cognitive science and psychology can complement AI evaluation research by providing methodological and interpretive frameworks that reveal new insights into model behavior.

大模型偏见推理公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。