arXiv:2511.19131cs.CL2025-11AAAI被引 2

通过优化隐藏状态激发基础大模型的思维链推理能力

Eliciting Chain-of-Thought in Base LLMs via Gradient-Based Representation Optimization

  • 将推理引导建模为带正则化的概率生成优化问题
  • 在数学、常识和逻辑推理任务上显著优于现有方法
  • 适合希望提升基础模型推理能力的研究者

思维链(Chain-of-Thought, CoT)推理是大语言模型处理复杂多步任务的关键能力。尽管基础模型在通用语料上预训练,但因缺乏专门训练而常难以进行推理,近期研究发现其隐藏状态中潜藏着推理潜能。然而,现有隐藏状态操纵方法(如线性激活引导)因结构僵硬且无约束,易导致分布偏移和文本质量下降。本文提出一种基于概率条件生成的新方法,将隐藏状态引导重构为推理导向轨迹,同时保持语言连贯性。通过平衡似然与先验正则化的优化框架,在数学、常识和逻辑推理多个基准测试中,本方法持续优于现有引导技术,为增强基础大模型的推理能力提供了理论严谨且高效的新路径。

原文摘要 · Abstract (English)

Chain-of-Thought (CoT) reasoning is a critical capability for large language models (LLMs), enabling them to tackle com- plex multi-step tasks. While base LLMs, pre-trained on general text corpora, often struggle with reasoning due to a lack of specialized training, recent studies reveal their latent reason- ing potential tied to hidden states. However, existing hidden state manipulation methods, such as linear activation steering, suffer from limitations due to their rigid and unconstrained nature, often leading to distribution shifts and degraded text quality. In this work, we propose a novel approach for elic- iting CoT reasoning from base LLMs through hidden state manipulation grounded in probabilistic conditional generation. By reformulating the challenge as an optimization problem with a balanced likelihood and prior regularization framework, our method guides hidden states toward reasoning-oriented trajectories while preserving linguistic coherence. Extensive evaluations across mathematical, commonsense, and logical reasoning benchmarks demonstrate that our approach con- sistently outperforms existing steering methods, offering a theoretically principled and effective solution for enhancing reasoning capabilities in base LLMs.

思维链推理增强隐藏状态优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。