通过干预模型推理过程,实现对大模型行为的精准控制。
Effectively Controlling Reasoning Models through Thinking Intervention
- 在推理过程中插入或修改关键思维标记,引导模型思考路径。
- 指令遵循任务准确率提升6.7%,安全拒绝率提高40.0%。
- 适合需要高可控性、安全性的复杂任务应用者使用。
增强推理能力的大语言模型(LLMs)会在生成最终答案前显式输出中间推理步骤,从而在复杂问题求解中表现优异。本文表明,这一新兴框架为更细粒度地控制模型行为提供了独特机会。我们提出「思维干预」(Thinking Intervention)新范式,通过战略性插入或修改特定思维标记,显式引导模型内部推理过程。实验发现,该范式在多种任务中显著提升推理模型性能:在IFEval和Overthinking上的指令遵循能力提升;在SEP上改善指令层级理解;在XSTest和SorryBench上增强安全对齐。结果表明,思维干预相比基线提示方法有显著优势,于开源DeepSeek R1模型上实现指令遵循准确率最高6.7%的提升,指令层级推理改善15.4%,对不安全提示的拒绝率提升40.0%。本工作开辟了控制推理型大模型的新研究方向。
原文摘要 · Abstract (English)
Reasoning-enhanced large language models (LLMs) explicitly generate intermediate reasoning steps prior to generating final answers, helping the model excel in complex problem-solving. In this paper, we demonstrate that this emerging generation framework offers a unique opportunity for more fine-grained control over model behavior. We propose Thinking Intervention, a novel paradigm designed to explicitly guide the internal reasoning processes of LLMs by strategically inserting or revising specific thinking tokens. We find that the Thinking Intervention paradigm enhances the capabilities of reasoning models across a wide range of tasks, including instruction following on IFEval and Overthinking, instruction hierarchy on SEP, and safety alignment on XSTest and SorryBench. Our results demonstrate that Thinking Intervention significantly outperforms baseline prompting approaches, achieving up to 6.7% accuracy gains in instruction-following scenarios, 15.4% improvements in reasoning about instruction hierarchies, and a 40.0% increase in refusal rates for unsafe prompts using open-source DeepSeek R1 models. Overall, our work opens a promising new research avenue for controlling reasoning LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。