arXiv:2508.10057q-bio.NCcs.AI2025-08被引 4

大模型在抽象推理中表现出与人类神经认知相似的特征。

Large Language Models Show Signs of Alignment with Human Neurocognition During Abstract Reasoning

  • 用脑电图对比人类与8个开源大模型的抽象模式识别表现。
  • 参数量达700亿的模型(如Qwen-2.5-72B)达到人类准确率,且难度感知一致。
  • 模型中间层对抽象模式的表征结构与人类前额叶脑电波相关,具潜在共享机制。

本研究探讨大语言模型(LLMs)在抽象推理过程中是否反映人类神经认知。我们对比了人类受试者与8个开源LLM在抽象模式补全任务中的表现及神经表征。通过分析任务表现和脑电图(EEG)记录的注视相关电位(FRPs)差异,发现仅最大规模的模型(约700亿参数)达到人类可比准确率,其中Qwen-2.5-72B与DeepSeek-R1-70B还展现出与人类一致的模式特定难度分布。所有模型在中间层均形成显著聚类的抽象模式类别表征,且聚类强度随任务性能提升。任务最优层的表征几何与人类前额叶FRPs呈中等正相关,而与其他脑电信号(响应锁定ERP、静息态EEG)无显著关联,暗示抽象模式可能存在共享表征空间。结果表明,大模型可能在抽象推理中映射人类脑机制,为生物与人工智能间存在共性原理提供初步证据。

原文摘要 · Abstract (English)

This study investigates whether large language models (LLMs) mirror human neurocognition during abstract reasoning. We compared the performance and neural representations of human participants with those of eight open-source LLMs on an abstract-pattern-completion task. We leveraged pattern type differences in task performance and in fixation-related potentials (FRPs) as recorded by electroencephalography (EEG) during the task. Our findings indicate that only the largest tested LLMs (~70 billion parameters) achieve human-comparable accuracy, with Qwen-2.5-72B and DeepSeek-R1-70B also showing similarities with the human pattern-specific difficulty profile. Critically, every LLM tested forms representations that distinctly cluster the abstract pattern categories within their intermediate layers, although the strength of this clustering scales with their performance on the task. Moderate positive correlations were observed between the representational geometries of task-optimal LLM layers and human frontal FRPs. These results consistently diverged from comparisons with other EEG measures (response-locked ERPs and resting EEG), suggesting a potential shared representational space for abstract patterns. This indicates that LLMs might mirror human brain mechanisms in abstract reasoning, offering preliminary evidence of shared principles between biological and artificial intelligence.

大模型神经认知抽象推理脑电图

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。