用大模型生成与人类相似的联想数据,研究语言模型中的隐性偏见。
The "LLM World of Words" English free association norms generated by large language models
- 用三款大模型对1.2万词进行联想生成,构建可比数据集
- 通过认知网络分析发现模型存在与社会一致的性别刻板印象
- 适合研究模型偏见、语义记忆结构的研究者使用
自由联想被广泛用于认知心理学和语言学,以研究概念知识的组织方式。近年来,利用类似方法探究大语言模型(LLM)中编码的知识成为新方向,尤其可用于揭示模型偏见。然而,缺乏与人类生成数据可比的大规模LLM自由联想规范数据,制约了该研究进展。为解决此问题,我们基于人类数据集“小世界词汇”(SWOW)的约12,000个提示词,使用Mistral、Llama3和Haiku三款大模型生成三个新的可比数据集,命名为“大模型世界词汇”(LWOW)。结合SWOW与LWOW数据,我们构建了人类与大模型的语义记忆认知网络模型,展示如何利用这些数据研究人类与大模型中的隐性偏见,如社会普遍存在的有害性别刻板印象。
原文摘要 · Abstract (English)
Free associations have been extensively used in cognitive psychology and linguistics for studying how conceptual knowledge is organized. Recently, the potential of applying a similar approach for investigating the knowledge encoded in LLMs has emerged, specifically as a method for investigating LLM biases. However, the absence of large-scale LLM-generated free association norms that are comparable with human-generated norms is an obstacle to this new research direction. To address this limitation, we create a new dataset of LLM-generated free association norms modeled after the "Small World of Words" (SWOW) human-generated norms consisting of approximately 12,000 cue words. We prompt three LLMs, namely Mistral, Llama3, and Haiku, with the same cues as those in the SWOW norms to generate three novel comparable datasets, the "LLM World of Words" (LWOW). Using both SWOW and LWOW norms, we construct cognitive network models of semantic memory that represent the conceptual knowledge possessed by humans and LLMs. We demonstrate how these datasets can be used for investigating implicit biases in humans and LLMs, such as the harmful gender stereotypes that are prevalent both in society and LLM outputs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。