通过模块社区分析大模型认知模式,揭示其与生物大脑的相似与差异。
Unraveling the cognitive patterns of Large Language Models through module communities
- 构建网络框架,将认知技能、模型架构与数据集关联。
- 发现大模型模块具分布式协同特征,依赖跨区域动态交互。
- 适合关注模型可解释性与高效微调策略的研究者。
大型语言模型(LLMs)在科学、工程和社会应用中带来深远影响,从科学发现到医疗诊断再到聊天机器人。尽管广泛应用,其内部机制仍隐藏于数十亿参数和复杂结构之中,难以理解。本文借鉴生物学中新兴认知的研究方法,提出一种基于网络的框架,连接认知技能、模型架构与数据集,推动基础模型分析范式革新。模块社区中的技能分布显示,虽然大模型未严格复制生物系统中特定功能的专一化,但其模块形成独特社区,其涌现的技能模式部分反映鸟类与小型哺乳动物大脑中分布式且相互关联的认知组织。数值结果表明,与生物系统相比,大模型在技能获取上更依赖动态的跨区域交互与神经可塑性。结合认知科学与机器学习,该框架为模型可解释性提供新视角,提示有效微调应利用分布式学习动态,而非僵化的模块干预。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have reshaped our world with significant advancements in science, engineering, and society through applications ranging from scientific discoveries and medical diagnostics to Chatbots. Despite their ubiquity and utility, the underlying mechanisms of LLM remain concealed within billions of parameters and complex structures, making their inner architecture and cognitive processes challenging to comprehend. We address this gap by adopting approaches to understanding emerging cognition in biology and developing a network-based framework that links cognitive skills, LLM architectures, and datasets, ushering in a paradigm shift in foundation model analysis. The skill distribution in the module communities demonstrates that while LLMs do not strictly parallel the focalized specialization observed in specific biological systems, they exhibit unique communities of modules whose emergent skill patterns partially mirror the distributed yet interconnected cognitive organization seen in avian and small mammalian brains. Our numerical results highlight a key divergence from biological systems to LLMs, where skill acquisition benefits substantially from dynamic, cross-regional interactions and neural plasticity. By integrating cognitive science principles with machine learning, our framework provides new insights into LLM interpretability and suggests that effective fine-tuning strategies should leverage distributed learning dynamics rather than rigid modular interventions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。