arXiv:2602.00986cs.CL2026-02

发现大模型隐藏层中存在稀疏奖励神经元,可预测模型信心与推理过程奖励。

Sparse Reward Subsystem in Large Language Models

  • 从大模型隐藏状态中识别出两类稀疏神经元:价值神经元和多巴胺神经元。
  • 价值神经元在不同数据集和模型间具有鲁棒性,且能准确预测模型置信度。
  • 该系统可用于推理时搜索引导,为理解模型决策提供新视角。

近期研究表明,大语言模型的隐藏状态编码了与奖励相关的信息,如答案正确性和模型置信度。然而,现有方法通常对全量隐藏状态进行黑箱探测,难以揭示信息在神经元中的组织结构。本文表明,奖励相关信息集中在一组稀疏神经元中。通过简单探测,我们识别出两类神经元:价值神经元(其激活预测状态价值),以及多巴胺神经元(其激活编码步骤级时间差分误差)。这两类神经元共同构成大模型隐藏状态中的稀疏奖励子系统。命名参照神经科学,因为生物奖励系统中的价值神经元与多巴胺神经元分别编码价值与奖励预测误差。我们证明价值神经元在多种数据集和模型间具有鲁棒性和可迁移性,并提供了因果证据表明其编码奖励相关信号。最后,我们展示了该子系统的应用:价值神经元可作为模型置信度的有效预测器,多巴胺神经元则可充当过程奖励模型(PRM),用于指导推理时的搜索。

原文摘要 · Abstract (English)

Recent studies show that LLM hidden states encode reward-related information, such as answer correctness and model confidence. However, existing approaches typically fit black-box probes on the full hidden states, offering little insight into how this information is structured across neurons. In this paper, we show that reward-related information is concentrated in a sparse subset of neurons. Using simple probing, we identify two types of neurons: value neurons, whose activations predict state value, and dopamine neurons, whose activations encode step-level temporal difference (TD) errors. Together, these neurons form a sparse reward subsystem within LLM hidden states. These names are drawn by analogy with neuroscience, where value neurons and dopamine neurons in the biological reward subsystem also encode value and reward prediction errors, respectively. We demonstrate that value neurons are robust and transferable across diverse datasets and models, and provide causal evidence that they encode reward-related information. Finally, we show applications of the reward subsystem: value neurons serve as effective predictors of model confidence, and dopamine neurons can function as a process reward model (PRM) to guide inference-time search.

大模型机制奖励建模神经元分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。