arXiv:2509.26507cs.NEcs.AI2025-09被引 6

受大脑启发的新型语言模型,兼具可解释性与Transformer性能。

The Dragon Hatchling: The Missing Link between the Transformer and Models of the Brain

  • 基于生物可实现的分形神经网络,模拟大脑结构与学习机制。
  • 在10M至10亿参数范围内,性能媲美GPT-2,且激活稀疏正向。
  • 支持单义性表达,适合追求可解释性的模型研究者使用。

计算系统与大脑的关系自冯·诺依曼和图灵时代便备受关注。生物网络具有无标度特性,能长期泛化,是机器学习迈向通用推理模型的主要障碍。本文提出‘龙雏’(BDH),一种基于n个局部交互神经元粒子的无标度生物启发型大语言模型架构。BDH兼具强理论基础与内在可解释性,性能媲美Transformer。它是一种基于注意力的状态空间序列学习架构,同时具备图模型属性和GPU友好设计。实证显示,在相同参数量(10M至10亿)与训练数据下,其在语言和翻译任务中表现接近GPT-2。BDH推理过程完全依赖突触可塑性与赫布学习的脉冲神经元。实验验证:处理特定概念时,个别突触会强化连接。其神经交互网络具有高模块性与重尾度分布。该模型生物合理,揭示了人类神经元实现语言的可能机制。模型激活向量稀疏且为正,展现单义性特征。状态可解释性是其固有特性,超越传统参数与神经元层面。

原文摘要 · Abstract (English)

The relationship between computing systems and the brain has served as motivation for pioneering theoreticians since John von Neumann and Alan Turing. Uniform, scale-free biological networks, such as the brain, have powerful properties, including generalizing over time, which is the main barrier for Machine Learning on the path to Universal Reasoning Models. We introduce `Dragon Hatchling' (BDH), a new Large Language Model architecture based on a scale-free biologically inspired network of \$n\$ locally-interacting neuron particles. BDH couples strong theoretical foundations and inherent interpretability without sacrificing Transformer-like performance. BDH is a practical, performant state-of-the-art attention-based state space sequence learning architecture. In addition to being a graph model, BDH admits a GPU-friendly formulation. It exhibits Transformer-like scaling laws: empirically BDH rivals GPT2 performance on language and translation tasks, at the same number of parameters (10M to 1B), for the same training data. BDH can be represented as a brain model. The working memory of BDH during inference entirely relies on synaptic plasticity with Hebbian learning using spiking neurons. We confirm empirically that specific, individual synapses strengthen connection whenever BDH hears or reasons about a specific concept while processing language inputs. The neuron interaction network of BDH is a graph of high modularity with heavy-tailed degree distribution. The BDH model is biologically plausible, explaining one possible mechanism which human neurons could use to achieve speech. BDH is designed for interpretability. Activation vectors of BDH are sparse and positive. We demonstrate monosemanticity in BDH on language tasks. Interpretability of state, which goes beyond interpretability of neurons and model parameters, is an inherent feature of the BDH architecture.

大模型生物启发可解释性神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。