提出四阶段框架,评估大模型能否作为认知科学解释模型。
On the use of foundation models in cognitive science

- 构建四阶段推理框架,连接模型行为与人类实验
- 强调仅行为对齐不足,需理论支撑和对比验证
- 适合关注认知建模与大模型交叉研究者
近期多项研究评估了基础模型(FMs)在认知与发育方面的对齐性,涵盖其在多种认知领域与成人表现的对应关系,以及模型训练过程是否反映儿童认知发展。然而,将FMs用作认知模型面临重大方法与概念挑战。核心问题在于:在何种条件下,行为对齐可支持将FMs视为认知的解释模型?本文提出一个四阶段推断框架:将人类实验任务适配为模型兼容格式、设定映射模型输出与人类测量的连接假设、评估行为对应性、在候选模型或干预之间进行比较。文章阐明连接假设的关键作用,识别限制对齐主张的挑战,并提出基于理论驱动与对比评估的原则。强调行为拟合本身不足以成立;只有嵌入明确理论承诺、理论诊断任务及系统性对比评估时,对齐才具科学意义。
原文摘要 · Abstract (English)
A host of recent studies have evaluated the cognitive and developmental alignment of Foundation Models (FMs). These investigations include evaluations of their correspondence to adult performance across a range of cognitive domains, as well as whether aspects of model training track children's cognitive development. However, using FMs as candidate cognitive models poses significant methodological and conceptual challenges. A key question underlies this effort: under what conditions does behavioral alignment justify treating FMs as explanatory models of cognition? In this paper, we articulate a four-stage inferential framework for evaluating FMs as cognitive and developmental models: adapting human experimental tasks to model-compatible formats, specifying linking hypotheses that map model outputs to human measures, evaluating behavioral correspondence, and comparing across candidate models or manipulations. We clarify the role of linking hypotheses in mapping model outputs to human behavioral measures, identify challenges that constrain alignment claims, and propose principles for theory-driven and comparative evaluation. Throughout, we argue that behavioral fit alone is insufficient. Alignment becomes scientifically meaningful only when embedded within explicit theoretical commitments, theory-diagnostic tasks, and systematic contrastive evaluation across candidate models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。