揭示大模型自我反思的内部激活轨迹,发现三阶段认知过程。
From Latent Signals to Reflection Behavior: Tracing Meta-Cognitive Activation Trajectory in R1-Style LLMs
- 通过分层追踪激活,发现思考预算、话语线索、行为输出三阶段机制。
- 反射行为令牌采样概率在最后层显著上升,与前序层调控直接相关。
- 适合研究大模型认知机制或可解释性的研究人员阅读。
R1风格大模型因其自我反思能力受到关注,但其内部机制仍不清晰。本文以反思行为的启动为锚点,追踪其逐层激活轨迹。利用logit lens读取标记级语义,发现有序进展:(i) 隐含控制层中,近似线性方向编码了思考预算语义;(ii) 语义枢纽层中,转折点与总结提示等话语级线索浮现并主导概率分布;(iii) 行为外显层中,反射行为令牌的生成概率逐渐升高,最终极易被采样。进一步干预实验揭示因果链:提示语义调节隐含控制方向上的激活投影,引发语义枢纽层中转折点与总结线索的竞争,进而调控行为外显层中反射令牌的采样概率。整体表明该过程类似人类元认知:从隐式监控,到话语级调节,再到显式反思。分析代码见 https://github.com/DYR1/S3-CoT。
原文摘要 · Abstract (English)
R1-style LLMs have attracted growing attention for their capacity for self-reflection, yet the internal mechanisms underlying such behavior remain unclear. To bridge this gap, we anchor on the onset of reflection behavior and trace its layer-wise activation trajectory. Using the logit lens to read out token-level semantics, we uncover a structured progression: (i) Latent-control layers, where an approximate linear direction encodes the semantics of thinking budget; (ii) Semantic-pivot layers, where discourse-level cues, including turning-point and summarization cues, surface and dominate the probability mass; and (iii) Behavior-overt layers, where the likelihood of reflection-behavior tokens begins to rise until they become highly likely to be sampled. Moreover, our targeted interventions uncover a causal chain across these stages: prompt-level semantics modulate the projection of activations along latent-control directions, thereby inducing competition between turning-point and summarization cues in semantic-pivot layers, which in turn regulates the sampling likelihood of reflection-behavior tokens in behavior-overt layers. Collectively, our findings suggest a human-like meta-cognitive process-progressing from latent monitoring, to discourse-level regulation, and to finally overt self-reflection. Our analysis code can be found at https://github.com/DYR1/S3-CoT.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。