发现并控制大模型的自我反思能力,提升推理效果与效率
From Emergence to Control: Probing and Modulating Self-Reflection in Language Models
- 通过注入反思线索激活预训练模型的隐藏反思能力
- 将反思频率从0.6%提升至18.6%,推理准确率最高提升12%
- 无需额外训练即可双向调控反思行为,适合需要高效推理的研究者
自反思——大语言模型重新审视、评估并修正自身推理的能力——最近在基于可验证奖励的强化学习(RLVR)微调后出现。尽管自反思与推理准确性提升相关,其起源和机制仍不明确。本文首次表明,自反思并非仅存在于RLVR微调模型中:它在预训练模型中也已出现,但极为罕见。为探测这一潜在能力,我们提出反射诱导探测法,将微调模型中的反思触发推理轨迹注入预训练模型。该干预使Qwen2.5的自反思频率从0.6%提升至18.6%,揭示其隐藏的反思潜能。进一步分析显示,预训练与微调模型均在内部表示中保留区分反思与非反思情境的隐状态。据此,我们构建了自反思向量——激活空间中与反思推理相关的一个方向。通过操纵该向量,我们实现了对预训练与微调模型自反思行为的双向控制。多推理基准实验表明,增强该向量可使推理性能最高提升12%,抑制则降低计算成本,提供无需额外训练即可在推理质量与效率间灵活权衡的机制。研究深化了对自反思的理解,并支持模型内部分析可实现精准行为控制的观点。
原文摘要 · Abstract (English)
Self-reflection -- the ability of a large language model (LLM) to revisit, evaluate, and revise its own reasoning -- has recently emerged as a powerful behavior enabled by reinforcement learning with verifiable rewards (RLVR). While self-reflection correlates with improved reasoning accuracy, its origin and underlying mechanisms remain poorly understood. In this work, {\it we first show that self-reflection is not exclusive to RLVR fine-tuned models: it already emerges, albeit rarely, in pretrained models}. To probe this latent ability, we introduce Reflection-Inducing Probing, a method that injects reflection-triggering reasoning traces from fine-tuned models into pretrained models. This intervention raises self-reflection frequency of Qwen2.5 from 0.6\% to 18.6\%, revealing a hidden capacity for reflection. Moreover, our analysis of internal representations shows that both pretrained and fine-tuned models maintain hidden states that distinctly separate self-reflective from non-reflective contexts. Leveraging this observation, {\it we then construct a self-reflection vector, a direction in activation space associated with self-reflective reasoning}. By manipulating this vector, we enable bidirectional control over the self-reflective behavior for both pretrained and fine-tuned models. Experiments across multiple reasoning benchmarks show that enhancing these vectors improves reasoning performance by up to 12\%, while suppressing them reduces computational cost, providing a flexible mechanism to navigate the trade-off between reasoning quality and efficiency without requiring additional training. Our findings further our understanding of self-reflection and support a growing body of work showing that understanding model internals can enable precise behavioral control.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。