arXiv:2603.18171cs.CL2026-03被引 1

对比大模型与人类的词汇联想,发现模型大小和温度影响联想多样性与典型性。

Modeling the human lexicon under temperature variations: linguistic factors, diversity and typicality in LLM word associations

  • 用不同温度下的大模型生成词对,对比人类与模型的联想模式。
  • 大模型更典型但变化少,小模型更多样但不够典型,高温提升多样性降低典型性。
  • 适合研究模型语言表征、认知模拟或评估生成质量的读者。

大型语言模型(LLMs)在文本生成流畅性上表现优异,但其内部词汇知识的人类相似性仍不明确。本研究通过比较人类与大模型生成的词联想对,评估模型对人类词汇模式的捕捉能力。基于SWOW数据集的英语提示-应答对,以及Mistral-7B、Llama-3.1-8B和Qwen-2.5-32B三款模型在多个温度设置下的新生成联想,分析了词汇频率、具体性等语言因素的影响,以及响应的可变性与典型性。结果显示,所有模型均反映人类在频率与具体性上的趋势,但在响应可变性和典型性上存在差异:如Qwen等大模型倾向于生成高度典型但变化小的“原型”响应,而较小模型如Mistral和Llama则产生更多样但较不典型的响应。温度越高,响应越多样,但典型性越低。这些发现揭示了人类与大模型词汇表征的异同,强调在探查模型词汇表示时需考虑模型规模与温度设置的影响。

原文摘要 · Abstract (English)

Large language models (LLMs) achieve impressive results in terms of fluency in text generation, yet the nature of their linguistic knowledge - in particular the human-likeness of their internal lexicon - remains uncertain. This study compares human and LLM-generated word associations to evaluate how accurately models capture human lexical patterns. Using English cue-response pairs from the SWOW dataset and newly generated associations from three LLMs (Mistral-7B, Llama-3.1-8B, and Qwen-2.5-32B) across multiple temperature settings, we examine (i) the influence of lexical factors such as word frequency and concreteness on cue-response pairs, and (ii) the variability and typicality of LLM responses compared to human responses. Results show that all models mirror human trends for frequency and concreteness but differ in response variability and typicality. Larger models such as Qwen tend to emulate a single "prototypical" human participant, generating highly typical but minimally variable responses, while smaller models such as Mistral and Llama produce more variable yet less typical responses. Temperature settings further influence this trade-off, with higher values increasing variability but decreasing typicality. These findings highlight both the similarities and differences between human and LLM lexicons, emphasizing the need to account for model size and temperature when probing LLM lexical representations.

词汇联想大模型认知温度影响语言表征

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。