小模型也能模拟人类认知,且能评估实验噪声上限
Small Foundation Models of Human Cognition and Behaviour

- 用135M到14B参数的模型在心理实验数据上训练,发现0.6B~1B参数已够用
- 小模型在新任务上泛化能力差,大模型在新结构任务中明显占优
- 模型依赖刺激和反馈信息,而非仅靠选择历史,适合心理学实验分析
在1070万条来自160项实验的试次级选择数据(Psych-101)上,我们训练了14个参数量从135M到14B、涵盖四种架构的模型。在分布内模拟中,模型性能几乎不随规模变化,0.6B至1B参数即可达到70B基线水平;而在分布外,大模型在新任务结构上的泛化能力显著更强。通过逐步屏蔽任务指令、实验刺激、结果反馈和选择历史四个提示通道,并打乱试次顺序,发现移除刺激和反馈内容会损失75.7%的已学信息,使模型表现低于随机水平,说明其依赖真实刺激与反馈,而非仅依赖选择历史。这表明小规模认知微调模型可作为心理实验的噪声上限估计器,但其应用范围仍受限于训练中见过的范式。
原文摘要 · Abstract (English)
Large language models fine-tuned on human behavioural data have emerged as general-purpose cognitive proxies, but the scale this requires, and whether these models process task structure or exploit statistical shortcuts, remain open questions. We train fourteen models from 135M to 14B parameters across four architecture families on Psych-101, a dataset of 10.7 million trial-level choices from 160 experiments. For in-distribution simulations, scale barely matters. The models fall within a narrow band, as though against a ceiling, and 0.6B to 1B parameters suffice to match a 70B baseline on held-out participants. Out-of-distribution, that band opens into a markedly steeper scaling gradient, with larger models clearly advantaged in generalisation to novel task structure. To determine what information these models use, we run two diagnostics. We progressively strip four prompt channels -- task instructions, experimental stimuli, outcome feedback, and choice history -- across 27 experiments, and permute trial order. Masking the content of stimuli and feedback destroys 75.7% of learned information and pushes models below chance, demonstrating that choice history alone does not account for performance. Permutation reveals invariance on tasks with independent trials but sensitivity where trial order is determined by prior responses. Small cognitively fine-tuned models therefore show promise as noise ceiling estimators for psychological experiments, though their scope remains bounded by the paradigms seen in training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。