发现大模型生成记忆时有类人策略,可提升人机协作效率
Emerging Human-like Strategies for Semantic Memory Foraging in Large Language Models
- 用语义流畅性任务分析大模型的生成搜索模式
- 模型在不同层中表现出收敛与发散的双重行为特征
- 为大模型认知对齐或互补增强提供新思路
人类和大型语言模型(LLMs)均存储大量语义记忆。在人类中,高效且有策略地访问这些记忆是多种认知功能的基础,这一机制长期受到心理学和计算科学的关注。其中,语义流畅性任务(Semantic Fluency Task, SFT)是一种广泛使用的神经心理评估工具,要求参与者尽可能多地生成具有语义约束的概念。本研究采用机制可解释性技术,以SFT为案例,深入探讨大模型中的语义记忆搜寻行为。重点分析了收敛与发散两种生成搜索模式,它们在人类中分别承担互补的战略角色,有助于高效记忆搜寻。我们发现,这些关键行为特征同样能在大模型的不同层级中被识别出来。该分析或可为大模型向人类认知更紧密对齐提供新视角,或引导其走向有益的认知非对齐,从而在人机交互中发挥互补优势。
原文摘要 · Abstract (English)
Both humans and Large Language Models (LLMs) store a vast repository of semantic memories. In humans, efficient and strategic access to this memory store is a critical foundation for a variety of cognitive functions. Such access has long been a focus of psychology and the computational mechanisms behind it are now well characterized. Much of this understanding has been gleaned from a widely-used neuropsychological and cognitive science assessment called the Semantic Fluency Task (SFT), which requires the generation of as many semantically constrained concepts as possible. Our goal is to apply mechanistic interpretability techniques to bring greater rigor to the study of semantic memory foraging in LLMs. To this end, we present preliminary results examining SFT as a case study. A central focus is on convergent and divergent patterns of generative memory search, which in humans play complementary strategic roles in efficient memory foraging. We show that these same behavioral signatures, critical to human performance on the SFT, also emerge as identifiable patterns in LLMs across distinct layers. Potentially, this analysis provides new insights into how LLMs may be adapted into closer cognitive alignment with humans, or alternatively, guided toward productive cognitive \emph{disalignment} to enhance complementary strengths in human-AI interaction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。