发现大模型认知模式与人类有相似也有差异,尤其在推理和创意上表现不一。
Humanlike Cognitive Patterns as Emergent Phenomena in Large Language Models
- 用心理学实验对比大模型与人类在决策、推理、创意上的表现
- 大模型具备类似人类的决策偏见和系统2式推理能力,但创意依赖语言生成
- 适合研究认知机制或人机协作的学者参考
大型语言模型(LLMs)中涌现的认知模式在心理学与人工智能领域引发广泛关注,亟需系统性综述以整合复杂的研究进展。本文系统回顾了大模型在决策偏差、推理和创造力三个关键认知领域的表现,基于心理学经典测试进行实证分析,并与人类基准对比。在决策方面,大模型展现出多种类人偏见,但部分人类特有偏见未出现,表明其认知模式仅部分与人类对齐;在推理方面,如GPT-4等先进模型表现出类似人类系统2的深思熟虑推理,而小模型则未达人类水平;在创造力方面,大模型在语言类创作任务(如故事生成)中表现优异,但在需要真实世界背景的发散思维任务中表现薄弱。尽管如此,研究提示大模型在人机协同问题解决中具有显著辅助潜力。同时,本文也指出记忆、注意力及开源模型发展等方面的局限,并为未来研究提供方向。
原文摘要 · Abstract (English)
Research on emergent patterns in Large Language Models (LLMs) has gained significant traction in both psychology and artificial intelligence, motivating the need for a comprehensive review that offers a synthesis of this complex landscape. In this article, we systematically review LLMs' capabilities across three important cognitive domains: decision-making biases, reasoning, and creativity. We use empirical studies drawing on established psychological tests and compare LLMs' performance to human benchmarks. On decision-making, our synthesis reveals that while LLMs demonstrate several human-like biases, some biases observed in humans are absent, indicating cognitive patterns that only partially align with human decision-making. On reasoning, advanced LLMs like GPT-4 exhibit deliberative reasoning akin to human System-2 thinking, while smaller models fall short of human-level performance. A distinct dichotomy emerges in creativity: while LLMs excel in language-based creative tasks, such as storytelling, they struggle with divergent thinking tasks that require real-world context. Nonetheless, studies suggest that LLMs hold considerable potential as collaborators, augmenting creativity in human-machine problem-solving settings. Discussing key limitations, we also offer guidance for future research in areas such as memory, attention, and open-source model development.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。