测试大模型能否模拟人类发音流畅性任务中的个体差异
Can LLMs Simulate Human Behavioral Variability? A Case Study in the Phonemic Fluency Task
- 对比34个模型在45种配置下的输出与106人数据
- 所有模型多样性均低于人类,新模型反而更单一
- 多模型集成也无法提升多样性,因词汇重叠高
大语言模型(LLMs)正被探索作为认知任务中人类参与者的替代品,但其模拟人类行为变异的能力尚不明确。本研究考察了LLMs在发音流畅性任务中是否能逼近个体差异——参与者需生成以特定字母开头的单词。我们评估了来自主流闭源与开源提供商的34个不同模型,在45种配置下输出,并与106名人类参与者的反应进行比较。尽管部分模型(尤其是Claude 3.7 Sonnet)近似了人类平均值和词汇偏好,但无一能再现人类的多样性范围。LLM输出始终缺乏多样性,且较新模型及启用思考模式的版本往往进一步降低多样性。网络分析揭示了人类与最接近人类的模型在检索结构上的根本差异。通过融合多个模型输出的集成模拟也未能恢复人类级别的多样性,可能源于各模型间存在高词汇重叠。这些结果凸显了使用LLMs模拟人类认知与行为的关键局限。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly explored as substitutes for human participants in cognitive tasks, but their ability to simulate human behavioral variability remains unclear. This study examines whether LLMs can approximate individual differences in the phonemic fluency task, where participants generate words beginning with a target letter. We evaluated 34 distinct models across 45 configurations from major closed-source and open-source providers, and compared outputs to responses from 106 human participants. While some models, especially Claude 3.7 Sonnet, approximated human averages and lexical preferences, none reproduced the scope of human variability. LLM outputs were consistently less diverse, with newer models and thinking-enabled modes often reducing rather than increasing variability. Network analysis further revealed fundamental differences in retrieval structure between humans and the most human-like model. Ensemble simulations combining outputs from diverse models also failed to recover human-level diversity, likely due to high vocabulary overlap across models. These results highlight key limitations in using LLMs to simulate human cognition and behavior.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。