arXiv:2604.12816cs.CL2026-04被引 1

揭示人类与大模型隐性偏见的认知差异,发现人类概念记忆结构可抑制偏见。

The role of System 1 and System 2 semantic memory structure in human and LLM biases

论文配图:The role of System 1 and System 2 semantic memory structure in human and LLM biases
图 1 · 摘自论文原文
  • 将人类与大模型的语义记忆建模为两种不同结构网络,模拟系统1与系统2思维。
  • 仅人类的系统2记忆结构显著降低隐性性别偏见,大模型无此效应。
  • 证明大模型缺乏人类特有的概念知识,难以自我调节偏见,适合认知科学与AI伦理研究者。

人类与大语言模型(LLMs)中的隐性偏见带来重大社会风险。双过程理论认为,偏见主要源于联想性的系统1思维,而反思性的系统2思维可缓解偏见,但其认知机制仍不明确。为更深入理解这一二元性在人类中的根源,可能也适用于大模型,我们基于人类与大模型生成的同类数据集,构建了具有不同结构的系统1与系统2语义记忆网络。通过基于网络的评估指标,探究这些结构与隐性性别偏见的关系。结果表明,语义记忆结构在人类中不可约简,而大模型则不具备此特性,说明大模型缺少某些类人概念知识。此外,只有在人类中,语义记忆结构与隐性偏见呈一致关联:系统2结构的偏见水平更低。这表明特定类型的概念知识有助于人类调节偏见,但对大模型无效,凸显了人机认知的根本差异。

原文摘要 · Abstract (English)

Implicit biases in both humans and large language models (LLMs) pose significant societal risks. Dual process theories propose that biases arise primarily from associative System 1 thinking, while deliberative System 2 thinking mitigates bias, but the cognitive mechanisms that give rise to this phenomenon remain poorly understood. To better understand what underlies this duality in humans, and possibly in LLMs, we model System 1 and System 2 thinking as semantic memory networks with distinct structures, built from comparable datasets generated by both humans and LLMs. We then investigate how these distinct semantic memory structures relate to implicit gender bias using network-based evaluation metrics. We find that semantic memory structures are irreducible only in humans, suggesting that LLMs lack certain types of human-like conceptual knowledge. Moreover, semantic memory structure relates consistently to implicit bias only in humans, with lower levels of bias in System~2 structures. These findings suggest that certain types of conceptual knowledge contribute to bias regulation in humans, but not in LLMs, highlighting fundamental differences between human and machine cognition.

认知科学大模型偏见语义记忆双过程理论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。