大模型能自发形成人类对颜色的高效分类系统。
Evolution and compression in LLMs: On the emergence of human-aligned categorization
- 用信息瓶颈理论分析大模型的颜色命名,发现其分类效率随规模提升。
- 通过迭代学习模拟文化演化,部分模型逼近人类近优分类效率。
- 仅具备强上下文学习能力的模型能复现人类多样化的分类模式,适合研究认知演化。
现有证据表明,人类语义分类体系通过信息瓶颈(IB)的复杂性-准确性权衡实现近似最优压缩。大语言模型(LLMs)并非为此目标训练,这引发疑问:它们能否演化出高效的、与人类对齐的语义系统?为此,我们聚焦颜色分类——一种具有丰富人类数据支持的认知理论测试基准,复现了两项经典人类研究。首先,在英语颜色命名实验中,我们发现不同大小的指令微调模型在复杂性和英语对齐度上差异显著,更大模型表现出更高对齐度和IB效率。其次,为检验这些模型是否仅模仿训练数据或真正具备类似人类的归纳偏好,我们提出迭代式上下文语言学习(IICLL)方法,模拟伪颜色命名系统的文化演化。结果发现,与人类相似,大模型会逐步将初始随机系统重构为更高IB效率的形式。然而,只有具备最强上下文学习能力的Gemini 2.0模型能重现人类观察到的广泛近优IB权衡范围,其他主流模型则收敛于低复杂性解。这些发现表明,人类对齐的语义分类可通过与人类相同的原理在大模型中自然涌现。
原文摘要 · Abstract (English)
Converging evidence suggests that human systems of semantic categories achieve near-optimal compression via the Information Bottleneck (IB) complexity-accuracy tradeoff. Large language models (LLMs) are not trained for this objective, which raises the question: are LLMs capable of evolving efficient human-aligned semantic systems? To address this question, we focus on color categorization -- a key testbed of cognitive theories of categorization with uniquely rich human data -- and replicate with LLMs two influential human studies. First, we conduct an English color-naming study, showing that LLMs vary widely in their complexity and English-alignment, with larger instruction-tuned models achieving better alignment and IB-efficiency. Second, to test whether these LLMs simply mimic patterns in their training data or actually exhibit a human-like inductive bias toward IB-efficiency, we simulate cultural evolution of pseudo color-naming systems in LLMs via a method we refer to as Iterated in-Context Language Learning (IICLL). We find that akin to humans, LLMs iteratively restructure initially random systems towards greater IB-efficiency. However, only a model with strongest in-context capabilities (Gemini 2.0) is able to recapitulate the wide range of near-optimal IB-tradeoffs observed in humans, while other state-of-the-art models converge to low-complexity solutions. These findings demonstrate how human-aligned semantic categories can emerge in LLMs via the same fundamental principle that underlies semantic efficiency in humans.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。