通过激活空间探索大模型反思机制,实现对反思行为的精准控制。
Unveiling the Latent Directions of Reflection in Large Language Models
- 用激活导向方法识别不同反思意图的潜在方向
- 可直接增强或抑制模型的反思行为,抑制更易实现
- 为防御恶意攻击和提升推理能力提供新思路
反思是大语言模型评估并修正自身推理的能力,广泛用于提升复杂推理任务的表现。然而,以往研究多聚焦于设计反思提示策略或强化学习目标,忽视了反思内在机制。本文从模型激活的潜在方向视角切入,提出基于激活引导的方法,刻画无反思、内在反思与触发反思三种不同反思意图。通过构建不同反思层级间的引导向量,实验表明:(1)可系统性识别新反射指令;(2)可通过激活干预直接增强或抑制反思行为;(3)抑制反思远比激发反思更容易。在GSM8k-adv与Cruxeval-o-adv数据集上,使用Qwen2.5-3B与Gemma3-4B-IT模型的实验显示各反思层级间存在清晰分层,且引导干预验证了反思行为的可控性。研究揭示了反思增强防御的潜力与对抗攻击中反思被抑制的风险,为理解大模型反思推理机制开辟新路径。
原文摘要 · Abstract (English)
Reflection, the ability of large language models (LLMs) to evaluate and revise their own reasoning, has been widely used to improve performance on complex reasoning tasks. Yet, most prior works emphasizes designing reflective prompting strategies or reinforcement learning objectives, leaving the inner mechanisms of reflection underexplored. In this paper, we investigate reflection through the lens of latent directions in model activations. We propose a methodology based on activation steering to characterize how instructions with different reflective intentions: no reflection, intrinsic reflection, and triggered reflection. By constructing steering vectors between these reflection levels, we demonstrate that (1) new reflection-inducing instructions can be systematically identified, (2) reflective behavior can be directly enhanced or suppressed through activation interventions, and (3) suppressing reflection is considerably easier than stimulating it. Experiments on GSM8k-adv and Cruxeval-o-adv with Qwen2.5-3B and Gemma3-4B-IT reveal clear stratification across reflection levels, and steering interventions confirm the controllability of reflection. Our findings highlight both opportunities (e.g., reflection-enhancing defenses) and risks (e.g., adversarial inhibition of reflection in jailbreak attacks). This work opens a path toward mechanistic understanding of reflective reasoning in LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。