语言模型通过文本预测自发形成类人概念表征。
Revealing emergent human-like conceptual representations from language prediction
- 在上下文推理中,模型从语言描述中灵活推导概念。
- 表征收敛到共享的无上下文结构,且与性能强相关。
- 结果与人类行为和脑活动高度一致,具生物合理性。
人类通过丰富的物理与社会经验习得概念,并用其理解与导航世界。相比之下,大型语言模型(LLMs)仅通过文本的下一个词预测训练,却表现出令人惊叹的人类行为特征。这些模型是否发展出类似人类的概念?若如此,这些概念如何表征、组织并关联行为?我们通过研究模型在上下文概念推断任务中形成的表征,回答了这些问题。发现模型能根据上下文线索,从语言描述中灵活推导概念。推导出的表征逐渐收敛至一个共享的、与上下文无关的结构,该结构的对齐程度可可靠预测模型在多种理解和推理任务中的表现。此外,这些收敛表征有效捕捉了人类行为判断,并与人脑神经活动模式高度吻合,提供了生物合理性的证据。综合来看,这些发现表明:仅通过语言预测,无需真实世界接地,也能自发产生结构化的人类类概念表征,凸显了概念结构在智能行为理解中的关键作用。更广泛而言,本工作表明,LLMs为探究人类概念本质提供了切实窗口,为推进人工智能与人类智能的对齐奠定了基础。
原文摘要 · Abstract (English)
People acquire concepts through rich physical and social experiences and use them to understand and navigate the world. In contrast, large language models (LLMs), trained solely through next-token prediction on text, exhibit strikingly human-like behaviors. Are these models developing concepts akin to those of humans? If so, how are such concepts represented, organized, and related to behavior? Here, we address these questions by investigating the representations formed by LLMs during an in-context concept inference task. We found that LLMs can flexibly derive concepts from linguistic descriptions in relation to contextual cues about other concepts. The derived representations converge toward a shared, context-independent structure, and alignment with this structure reliably predicts model performance across various understanding and reasoning tasks. Moreover, the convergent representations effectively capture human behavioral judgments and closely align with neural activity patterns in the human brain, providing evidence for biological plausibility. Together, these findings establish that structured, human-like conceptual representations can emerge purely from language prediction without real-world grounding, highlighting the role of conceptual structure in understanding intelligent behavior. More broadly, our work suggests that LLMs offer a tangible window into the nature of human concepts and lays the groundwork for advancing alignment between artificial and human intelligence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。