arXiv:2505.12822cs.AI2025-05被引 2

发现语言模型中专精罕见词的神经元,揭示其动态演化机制。

Emergent Specialization: Rare Token Neurons in Language Models

  • 识别出影响罕见词预测的关键神经元,具三段式动态组织结构。
  • 这些神经元形成协同激活子网络,避免与其他神经元共激活。
  • 机制或与重尾权重分布相关,适合研究模型内部表征者阅读。

大语言模型在表示和生成罕见词方面存在困难,尽管这些词在专业领域中至关重要。本研究识别出对语言模型罕见词预测具有异常强影响力的神经元结构,称为稀有标记神经元,并探究其涌现机制与行为特征。这些神经元在训练过程中动态演化出典型的三阶段组织结构(平台期、幂律期与快速衰减期),由初始同质状态发展为功能分化架构。在激活空间中,稀有标记神经元构成一个协调子网络,选择性地共同激活,同时避免与其他神经元共激活。这种功能特化可能与重尾权重分布的发展相关,暗示其背后存在统计力学基础。

原文摘要 · Abstract (English)

Large language models struggle with representing and generating rare tokens despite their importance in specialized domains. In this study, we identify neuron structures with exceptionally strong influence on language model's prediction of rare tokens, termed as rare token neurons, and investigate the mechanism for their emergence and behavior. These neurons exhibit a characteristic three-phase organization (plateau, power-law, and rapid decay) that emerges dynamically during training, evolving from a homogeneous initial state to a functionally differentiated architecture. In the activation space, rare token neurons form a coordinated subnetwork that selectively co-activates while avoiding co-activation with other neurons. This functional specialization potentially correlates with the development of heavy-tailed weight distributions, suggesting a statistical mechanical basis for emergent specialization.

语言模型神经元机制罕见词表征学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。