用真实科研轨迹测试大模型能否真正模拟人类认知。
Can Large Language Models Simulate Human Cognition Beyond Behavioral Imitation?
- 基于217位研究者长期科研记录构建认知评估基准
- 大模型在跨领域时间迁移任务中表现有限,难以复制个体认知模式
- 提出多维对齐指标,适合评估认知一致性而非表面模仿
人工智能的核心问题之一是大语言模型能否真正模拟人类认知,还是仅停留在行为表面的模仿。现有数据集或依赖合成推理路径,或仅提供群体层面聚合,无法捕捉个体真实认知轨迹。本文基于217位不同领域人工智能研究者的长期科研发表记录,将学术成果作为其认知过程的外部表征,构建了一个新的评估基准。为区分模型是迁移认知模式还是单纯模仿行为,基准采用跨领域、时间偏移的泛化设置。进一步提出多维认知对齐度量,用于评估个体层面的认知一致性。通过对主流大模型及多种增强技术的系统评估,首次实证回答两个关键问题:(1)当前大模型在模拟人类认知方面表现如何?(2)现有技术手段能将其能力提升到何种程度?
原文摘要 · Abstract (English)
An essential problem in artificial intelligence is whether LLMs can simulate human cognition or merely imitate surface-level behaviors, while existing datasets suffer from either synthetic reasoning traces or population-level aggregation, failing to capture authentic individual cognitive patterns. We introduce a benchmark grounded in the longitudinal research trajectories of 217 researchers across diverse domains of artificial intelligence, where each author's scientific publications serve as an externalized representation of their cognitive processes. To distinguish whether LLMs transfer cognitive patterns or merely imitate behaviors, our benchmark deliberately employs a cross-domain, temporal-shift generalization setting. A multidimensional cognitive alignment metric is further proposed to assess individual-level cognitive consistency. Through systematic evaluation of state-of-the-art LLMs and various enhancement techniques, we provide a first-stage empirical study on the questions: (1) How well do current LLMs simulate human cognition? and (2) How far can existing techniques enhance these capabilities?
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。