用认知神经科学方法解析大模型内部功能组织,揭示幻觉等缺陷的内在机制。
NeuroCogMap Reveals Cognitive Organization of Large Language Models

- 借鉴神经科学框架,将大模型内部特征划分为功能模块。
- 发现大模型失败对应特定表征与行为控制系统的异常破坏。
- 可帮助理解人类语言认知,适合研究模型可解释性与认知科学交叉者。
理解复杂认知功能在人工系统中的组织方式,是解读大语言模型(LLMs)并将其与生物认知关联的核心问题。尽管LLMs表现出广泛的类认知行为,但其内部表征是否形成可重复的功能系统以解释行为、失效及与人类认知的联系仍不明确。本文提出NeuroCogMap,一种受认知神经科学启发的框架,将LLM内部特征组织为功能区域,并将其与可解释的功能、认知能力及认知层级相连接。这些区域构成稳定且语义连贯的组织结构,在不同模型间部分保守,并与模型输出功能相关。在此结构中,大模型的主要失败模式——包括幻觉、偏见、拒绝失败和阿谀奉承——对应于表征与行为控制系统中的特定破坏,产生可指导检测与干预的内部签名。此外,NeuroCogMap提升了对人类自然语言理解过程中皮层反应的预测能力,尤其在高阶联合皮层中对应最强。在认知层面,其内部签名揭示了引导经典人类决策模型改进的潜在策略。这些发现确立了NeuroCogMap作为映射人工系统功能组织并将其与人类皮层功能及认知行为关联的系统级框架。
原文摘要 · Abstract (English)
Understanding how complex cognitive functions are organized within artificial systems is central to interpreting large language models (LLMs) and relating them to biological cognition. Yet although LLMs exhibit broad cognitive-like behaviours, it remains unclear whether their internal representations form reproducible functional systems that explain behaviour, failure and links to human cognition. Here we present NeuroCogMap, a cognitive neuroscience-inspired framework that organizes internal features of LLMs into functional parcels and links them to interpretable functions, cognitive capabilities and a cognitive hierarchy. These parcels form a stable and semantically coherent organization that is partly conserved across models and functionally linked to model outputs. Within this organization, major LLM failures, including hallucination, bias, refusal failure and sycophancy, correspond to distinct disruptions in representational and behavioural-control systems, yielding internal signatures for mechanism-guided detection and targeted intervention. Beyond model behaviour, NeuroCogMap improves prediction of human cortical responses during naturalistic language comprehension, with the strongest correspondence in higher-order association cortex. At the cognitive level, its internal signatures expose latent strategies that guide refinements of classical models of human decision-making. Together, these findings establish NeuroCogMap as a system-level framework for mapping functional organization in artificial systems and for relating this organization to human cortical function and cognitive behaviour.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。