arXiv:2606.14117stat.MEcs.AI2026-06

用心理测试方法评估大模型的隐性偏见,发现不同模型差异明显。

A Two-Stage Statistical Framework for Evaluating Associative Interference in Large Language Models

论文配图:A Two-Stage Statistical Framework for Evaluating Associative Interference in Large Language Models
图 1 · 摘自论文原文
  • 分两阶段建模,区分响应服从与任务一致性
  • Claude Sonnet-4在性别职业领域干扰达0.086,其他模型较弱
  • 揭示模型特性决定偏见程度,适合模型安全评估者

大型语言模型(LLMs)常通过人类心理学范式评估偏见,但方法局限——拒绝行为与任务表现混杂,影响解读。本文将内隐联想测验(IAT)改造为受控的强制选择框架,提出两阶段建模方法,分离响应合规性与任务一致性。在三个现代模型(Claude Sonnet-4、Gemini 2.5 Pro、GPT-5)上评估关联干扰,定义为不一致条件下的任务一致性下降。尽管响应合规性均高,干扰效应在模型和领域间差异显著:Claude Sonnet-4在性别-职业领域显示强干扰(DeltaP = 0.086,95% CrI [0.026, 0.173]),性别-科学领域亦有可信效应;Gemini 2.5 Pro干扰减弱,GPT-5则几乎无显著干扰。结果表明,IAT式关联不对称并非所有模型共有特征,而是取决于模型自身特性。通过分离干扰与合规性并建模个体项变异,本研究提供评估模型结构化响应模式的严谨框架,强调需按模型定制评估,并提示现代系统可显著缓解关联干扰。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly evaluated for bias using adaptations of human psychological paradigms, yet methodological limitations-particularly the conflation of refusal behavior with task performance-have hindered clear interpretation. Here, we adapt the Implicit Association Test (IAT) to a controlled, forced-choice framework and introduce a two-stage modeling approach that separates response compliance from task-consistent classification. Across three contemporary LLMs (Claude Sonnet-4, Gemini 2.5 Pro, and GPT-5), we evaluate associative interference, defined as reduced task-consistency in incongruent relative to congruent conditions. While compliance with the structured response format was uniformly high, interference effects varied substantially across models and domains. Claude Sonnet-4 exhibited strong interference in the Gender--Career domain (DeltaP = 0.086, 95% CrI [0.026, 0.173]) and smaller but credible effects in Gender--Science. Gemini 2.5 Pro showed attenuated interference, and GPT-5 exhibited minimal or no detectable interference across domains. These findings demonstrate that IAT-style associative asymmetries are not a universal property of LLMs, but instead depend on model-specific characteristics. By isolating interference from compliance and modeling item-level variability, this study provides a principled framework for evaluating structured response patterns in LLMs. The results highlight the importance of model-specific assessment and suggest that associative interference can be substantially mitigated in modern systems.

模型偏见心理测试大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。