arXiv:2506.18167cs.LGcs.AI2025-06中稿 · the Workshop on Re…被引 90

用向量控制大模型推理过程,让思考更可控可解释

Understanding Reasoning in Thinking Language Models via Steering Vectors

  • 通过分析激活空间中的线性方向,识别推理行为模式
  • 在500个任务上验证,可精准调节回溯与不确定表达
  • 适合需要可解释推理的AI系统开发者使用

近期大语言模型发展出思维型模型,可在生成回答前输出长链内部推理。尽管性能提升显著,但控制其推理过程仍具挑战。本文针对DeepSeek-R1-Distill模型,通过在10个不同类别共500个任务上的系统实验,识别出思维模型的多种推理行为,包括表达不确定性、生成示例验证假设、以及推理链中的回溯。我们发现这些行为由模型激活空间中的线性方向所调控,并可通过引导向量进行干预。通过提取并应用这些向量,可实现对模型推理倾向(如回溯频率或不确定表达)的可控调节。该方法在三个不同架构的DeepSeek-R1-Distill模型上得到验证,表现出一致的控制效果,为思维型模型提供了可解释且可控的推理引导工具。

原文摘要 · Abstract (English)

Recent advances in large language models (LLMs) have led to the development of thinking language models that generate extensive internal reasoning chains before producing responses. While these models achieve improved performance, controlling their reasoning processes remains challenging. This work presents a steering approach for thinking LLMs by analyzing and manipulating specific reasoning behaviors in DeepSeek-R1-Distill models. Through a systematic experiment on 500 tasks across 10 diverse categories, we identify several reasoning behaviors exhibited by thinking models, including expressing uncertainty, generating examples for hypothesis validation, and backtracking in reasoning chains. We demonstrate that these behaviors are mediated by linear directions in the model's activation space and can be controlled using steering vectors. By extracting and applying these vectors, we provide a method to modulate specific aspects of the model's reasoning process, such as its tendency to backtrack or express uncertainty. Our approach offers practical tools for steering reasoning processes in thinking models in a controlled and interpretable manner. We validate our steering method using three DeepSeek-R1-Distill models, demonstrating consistent control across different model architectures.

思维链推理控制可解释性向量引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。