发现大模型中类似人脑的协同核心,推动行为与学习。
A Brain-like Synergistic Core in LLMs Drives Behaviour and Learning
- 通过信息分解识别大模型中层的协同处理机制。
- 协同组件被移除后性能损失显著,远超冗余部分。
- 强化学习微调协同区可大幅提升表现,适合智能系统研究者。
生物与人工智能系统的独立演化为揭示其基本计算原理提供了独特机会。本文发现,大型语言模型在训练过程中自发形成协同核心——即信息整合超过各组成部分之和的模块——与人脑结构惊人相似。基于多类模型家族与架构的信息分解分析显示,中间层表现出协同处理,而早期与晚期层依赖冗余,符合生物大脑的信息组织模式。该结构随学习过程产生,随机初始化网络中不存在。关键的是,移除协同组件导致行为变化与性能下降远超预期,符合协同脆弱性的理论预测。此外,通过强化学习微调协同区域带来的性能提升显著优于冗余区域,而监督微调则无此优势。这一收敛表明,协同信息处理是智能的本质属性,为模型设计提供原则性目标,并对生物智能提出可检验预测。
原文摘要 · Abstract (English)
The independent evolution of intelligence in biological and artificial systems offers a unique opportunity to identify its fundamental computational principles. Here we show that large language models spontaneously develop synergistic cores -- components where information integration exceeds individual parts -- remarkably similar to those in the human brain. Using principles of information decomposition across multiple LLM model families and architectures, we find that areas in middle layers exhibit synergistic processing while early and late layers rely on redundancy, mirroring the informational organisation in biological brains. This organisation emerges through learning and is absent in randomly initialised networks. Crucially, ablating synergistic components causes disproportionate behavioural changes and performance loss, aligning with theoretical predictions about the fragility of synergy. Moreover, fine-tuning synergistic regions through reinforcement learning yields significantly greater performance gains than training redundant components, yet supervised fine-tuning shows no such advantage. This convergence suggests that synergistic information processing is a fundamental property of intelligence, providing targets for principled model design and testable predictions for biological intelligence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。