arXiv:2510.12680cs.LGcs.AI2025-10被引 8

研究大模型在思考与直接回答间的切换机制,发现现有方法存在行为泄露问题。

Demystifying Hybrid Thinking: Can LLMs Truly Switch Between Think and No-Think?

  • 通过四因素分析提升混合思维的可控性,包括数据规模和训练策略优化
  • 在MATH500上将非思考模式输出长度从1085降至585,推理提示词减少超90%
  • 适合关注大模型推理效率与可控性的研究人员及应用开发者

混合思维使大模型能在推理与直接回答间切换,兼顾效率与推理能力。但实验显示当前方法仅实现部分模式分离:推理行为常渗入非思考模式。我们分析影响可控性的关键因素,发现四大关键:(1) 更大的数据规模,(2) 使用不同问题的思考与非思考答案而非同一问题,(3) 适度增加非思考数据量,(4) 两阶段训练策略——先训练推理能力,再进行混合思维训练。基于此,提出实用训练方案:相比标准训练,在保持双模式准确率的同时,显著缩短非思考输出长度(MATH500上从1085降至585),并大幅减少推理提示词如"wait"的出现次数(从5917降至522)。研究揭示当前混合思维的局限,为增强其可控性提供方向。

原文摘要 · Abstract (English)

Hybrid thinking enables LLMs to switch between reasoning and direct answering, offering a balance between efficiency and reasoning capability. Yet our experiments reveal that current hybrid thinking LLMs only achieve partial mode separation: reasoning behaviors often leak into the no-think mode. To understand and mitigate this, we analyze the factors influencing controllability and identify four that matter most: (1) larger data scale, (2) using think and no-think answers from different questions rather than the same question, (3) a moderate increase in no-think data number, and (4) a two-phase strategy that first trains reasoning ability and then applies hybrid think training. Building on these findings, we propose a practical recipe that, compared to standard training, can maintain accuracy in both modes while significantly reducing no-think output length (from $1085$ to $585$ on MATH500) and occurrences of reasoning-supportive tokens such as ``\texttt{wait}'' (from $5917$ to $522$ on MATH500). Our findings highlight the limitations of current hybrid thinking and offer directions for strengthening its controllability.

大模型混合思维推理控制训练策略

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。