arXiv:2503.11572cs.CYcs.AI2025-03被引 3

发现推理模型处理反刻板印象信息时更费力,揭示其隐性偏见模式差异。

Implicit Bias-Like Patterns in Reasoning Models

  • 设计新测试法RM-IAT,探测推理模型的隐性偏见式计算过程。
  • o3-mini等模型在反刻板任务上多用约15%的推理令牌,表明计算负担更大。
  • 不同模型表现差异显著,反映其内部对偏见与刻板印象的处理机制不同。

隐性偏见指自动化的心理过程,影响感知、判断与行为。以往对大语言模型(LLMs)中‘隐性偏见’的研究主要关注输出结果,而非生成输出的内在过程。本文提出推理模型隐性关联测试(RM-IAT),用于研究采用逐步推理解决复杂任务的推理模型中的隐性偏见式处理模式。通过RM-IAT,我们发现o3-mini、DeepSeek-R1、gpt-oss-20b和Qwen-3 8B等模型在处理与既有联想不一致的任务时,比一致任务多消耗约15%的推理令牌,表明处理反刻板印象信息需要更多计算资源。相反,Claude 3.7 Sonnet表现出相反模式,主题分析将其归因于其内部对偏见与刻板印象的特别关注。这些发现表明,推理模型存在显著的隐性偏见式模式,且其表现受模型内部推理内容的影响而有明显差异。

原文摘要 · Abstract (English)

Implicit biases refer to automatic mental processes that shape perceptions, judgments, and behaviors. Previous research on "implicit bias" in LLMs focused primarily on outputs rather than the processes underlying the outputs. We present the Reasoning Model Implicit Association Test (RM-IAT) to study implicit bias-like processing in reasoning models, LLMs that use step-by-step reasoning to solve complex tasks. Using RM-IAT, we find that reasoning models like o3-mini, DeepSeek-R1, gpt-oss-20b, and Qwen-3 8B consistently expend more reasoning tokens on association-incompatible tasks than association-compatible tasks, suggesting greater computational effort when processing counter-stereotypical information. Conversely, Claude 3.7 Sonnet exhibited reversed patterns, which thematic analysis associated with its unique internal focus on reasoning about bias and stereotypes. These findings demonstrate that reasoning models exhibit distinct implicit bias-like patterns and that these patterns vary significantly depending on the models' internal reasoning content.

隐性偏见推理模型大模型机制认知偏差

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。