arXiv:2506.15617cs.CLcs.AI2025-06被引 3

揭示大模型表达后悔的内部机制,找到关键神经元与层级。

The Compositional Architecture of Regret in Large Language Models

  • 设计提示场景构建后悔数据集,定位模型输出中的后悔表达。
  • 提出S-CDI指标确定最优后悔表征层,识别出M型解耦模式。
  • 用RDS和GIC分析神经元功能,发现三类不同作用的神经元。

大语言模型中的后悔指其在面对与先前生成错误信息相矛盾的证据时,明确表达出的后悔情绪。研究这一机制对提升模型可靠性、揭示神经网络中的认知编码至关重要。为理解该机制,需先识别模型输出中的后悔表达,再分析其内部表征。这要求考察模型隐藏状态中神经元层面的信息处理过程。然而,当前面临三大挑战:(1) 缺乏专门捕捉后悔表达的数据集;(2) 无有效度量标准以确定最优后悔表征层;(3) 缺乏用于识别和分析后悔神经元的度量。针对上述问题,我们提出:(1) 通过精心设计的提示场景构建综合性后悔数据集的工作流程;(2) 提出监督压缩-解耦指数(S-CDI)以识别最优后悔表征层;(3) 提出后悔主导得分(RDS)识别后悔神经元,以及组影响系数(GIC)分析激活模式。实验结果表明,使用S-CDI成功定位了最优后悔表征层,显著提升了探测分类任务性能。此外,我们发现跨层存在M型解耦模式,揭示信息处理在耦合与解耦阶段间交替进行。通过RDS,我们将神经元分为三类:后悔神经元、非后悔神经元和双功能神经元。

原文摘要 · Abstract (English)

Regret in Large Language Models refers to their explicit regret expression when presented with evidence contradicting their previously generated misinformation. Studying the regret mechanism is crucial for enhancing model reliability and helps in revealing how cognition is coded in neural networks. To understand this mechanism, we need to first identify regret expressions in model outputs, then analyze their internal representation. This analysis requires examining the model's hidden states, where information processing occurs at the neuron level. However, this faces three key challenges: (1) the absence of specialized datasets capturing regret expressions, (2) the lack of metrics to find the optimal regret representation layer, and (3) the lack of metrics for identifying and analyzing regret neurons. Addressing these limitations, we propose: (1) a workflow for constructing a comprehensive regret dataset through strategically designed prompting scenarios, (2) the Supervised Compression-Decoupling Index (S-CDI) metric to identify optimal regret representation layers, and (3) the Regret Dominance Score (RDS) metric to identify regret neurons and the Group Impact Coefficient (GIC) to analyze activation patterns. Our experimental results successfully identified the optimal regret representation layer using the S-CDI metric, which significantly enhanced performance in probe classification experiments. Additionally, we discovered an M-shaped decoupling pattern across model layers, revealing how information processing alternates between coupling and decoupling phases. Through the RDS metric, we categorized neurons into three distinct functional groups: regret neurons, non-regret neurons, and dual neurons.

大模型后悔机制神经元分析表征学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。