arXiv:2503.02057cs.CLcs.AI2025-03

基于神经元局部学习机制,构建无需标注数据的语言模型。

Hebbian learning the local structure of language

  • 模拟大脑局部无监督学习,分层神经元自动分词
  • 无需训练即可生成符合真实语言形态分布的词汇结构
  • 适合研究语言起源与人类创造力的神经机制

大脑学习具有局部性和无监督性(海布式)。我们基于这一微观约束,推导出一种高效的人类语言模型。该模型包含两部分:(1) 分层神经元网络,从文本中学习词单元(即阅读时你做的事);(2) 额外神经元将已学语义模式组合为有语义的词元(嵌入表示)。模型支持持续并行学习且不遗忘;同时作为强大分词器,具备重整化群特性,可利用冗余信息,生成始终可分解为基集(如字母)的词元,并融合多语言特征。该模型结构使其能在无数据条件下学习自然语言形态,生成的语言预测出真实语言中词形模式的正确分布,进一步解释了为何人类语言被划分为词。该模型为理解语言与人类创造力的微观起源提供了基础。

原文摘要 · Abstract (English)

Learning in the brain is local and unsupervised (Hebbian). We derive the foundations of an effective human language model inspired by these microscopic constraints. It has two parts: (1) a hierarchy of neurons which learns to tokenize words from text (whichiswhatyoudowhenyoureadthis); and (2) additional neurons which bind the learned symanticless patterns of the tokenizer into a symanticful token (an embedding). The model permits continuous parallel learning without forgetting; and is a powerful tokenizer which performs renormalization group. This allows it to exploit redundancy, such that it generates tokens which are always decomposable into a basis set (e.g an alphabet), and can mix features learned from multiple languages. We find that the structure of this model allows it to learn a natural language morphology WITHOUT data. The language data generated by this model predicts the correct distribution of word-forming patterns observed in real languages, and further demonstrates why microscopically human speech is broken up into words. This model provides the basis for understanding the microscopic origins of language and human creativity.

语言建模海布学习无监督神经机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。