受大脑启发的新型语言模型,让AI更接近通用智能。
BriLLM: Brain-inspired Large Language Model
- 用神经信号流动机制替代传统Transformer结构。
- 1-20亿参数模型已实现类似GPT-1的生成能力。
- 适合研究脑科学与通用人工智能交叉方向的人看。
我们提出BriLLM,一种受大脑启发的大规模语言模型,通过实现信号全连接流动(SiFu)学习,从根本上重构机器学习基础。该工作解决了语言模型与‘世界模型’脱节、以及基于Transformer架构的根本性局限问题。BriLLM融合两项神经认知原则:(1) 静态语义映射,将词元映射到类皮层区域的专用节点;(2) 动态信号传播,模拟脑电活动中的信息动态。该架构实现多项突破:天然支持多模态、模型可解释性达节点级、上下文长度无关扩展、首次实现语言任务的全局脑样信息处理模拟。初步1-2B参数模型成功复现GPT-1级生成能力,并展现稳定困惑度下降。可扩展性分析证实100-200B参数版本可行,支持40,000词表规模。该范式符合奥卡姆剃刀原则(直接语义映射的简洁性),也契合自然演化规律(大脑经验证的通用智能架构)。BriLLM建立了一个生物根基扎实的通用智能新框架,突破现有方法根本瓶颈。
原文摘要 · Abstract (English)
We introduce BriLLM, a brain-inspired large language model that fundamentally redefines the foundations of machine learning through its implementation of Signal Fully-connected flowing (SiFu) learning. This work addresses the critical bottleneck hindering AI's progression toward Artificial General Intelligence (AGI)--the disconnect between language models and "world models"--as well as the fundamental limitations of Transformer-based architectures rooted in the conventional representation learning paradigm. BriLLM incorporates two pivotal neurocognitive principles: (1) static semantic mapping, where tokens are mapped to specialized nodes analogous to cortical areas, and (2) dynamic signal propagation, which simulates electrophysiological information dynamics observed in brain activity. This architecture enables multiple transformative breakthroughs: natural multi-modal compatibility, full model interpretability at the node level, context-length independent scaling, and the first global-scale simulation of brain-like information processing for language tasks. Our initial 1-2B parameter models successfully replicate GPT-1-level generative capabilities while demonstrating stable perplexity reduction. Scalability analyses confirm the feasibility of 100-200B parameter variants capable of processing 40,000-token vocabularies. The paradigm is reinforced by both Occam's Razor--evidenced in the simplicity of direct semantic mapping--and natural evolution--given the brain's empirically validated AGI architecture. BriLLM establishes a novel, biologically grounded framework for AGI advancement that addresses fundamental limitations of current approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。