arXiv:2510.01030cs.AI2025-10被引 1

发现让大模型像人一样理解概念的关键计算因素。

Uncovering the Computational Ingredients of Human-Like Representations in LLMs

  • 用认知科学的配对相似度任务评估75个模型的概念表征
  • 指令微调和注意力头维度是决定人类对齐度的核心因素
  • 现有评测基准无法全面反映模型与人的概念一致性

人类将多源感知与语言输入转化为结构化行为的能力,被认为依赖于对概念的稳健表征。尽管基于Transformer的大语言模型(LLMs)在架构、微调方法和训练数据等方面展现出多样化的计算要素,但哪些对构建类人概念表征最为关键仍不明确。此外,现有多数评测基准难以有效衡量表征对齐性,导致模型得分不可靠。本文通过在75个模型上使用认知科学中成熟的三元组相似度任务(基于THINGS数据库的概念),系统评估其表征对齐程度。结果表明,指令微调和更大的注意力头维度是预测人类对齐度最强的因素,而激活函数选择、多模态预训练和参数量影响较小。对齐分数与现有基准分数的相关性分析显示,尽管部分基准(如BigBenchHard)比其他(如MUSR)更能反映表征对齐,但均未能解释人类-模型对齐性的全部方差,说明其局限性。研究揭示了推动LLMs作为人类认知模型发展的关键计算要素,并填补了评估体系中的重要空白。

原文摘要 · Abstract (English)

The human ability to translate diverse perceptual and linguistic inputs into structured behavior has been thought to rest on learning robust representations of concepts. The rapid advancement of transformer-based large language models (LLMs) has surfaced a diversity of computational ingredients relevant for model building - architectures, fine-tuning methods, and training datasets among others - yet it remains unclear which are most crucial for developing human-like conceptual representations. Further, most current benchmarks are ill-suited to measuring representational alignment, making LLMs' scores on them unreliable for assessing whether they are progressing as cognitive models. We address these limitations by evaluating over 75 models on a triplet similarity task, a method well established in cognitive science for measuring conceptual representations, using concepts from the THINGS database. We find that instruction fine-tuning and larger attention head dimensionality are among the strongest predictors of human alignment, while activation function choice, multimodal pretraining, and parameter size have limited influence on alignment. Correlations between alignment scores and existing benchmark scores reveal that while some benchmarks (e.g., BigBenchHard) better capture representational alignment than others (e.g., MUSR), none fully accounts for the variance in human-model alignment, demonstrating their insufficiency. Taken together, our findings highlight key computational ingredients for advancing LLMs as models of human conceptual representation and address a key gap in LLM evaluation.

大模型概念表征认知评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。