思考模式让部分指令遵循任务变好,但另一些反而更差,关键看任务类型。
When Built-in Thinking Helps and Hurts: Constraint-Level Error Shifts in Instruction Following

- 通过开关思维模式对比,发现思考改变错误分布而非整体降分。
- 规划类任务因思考提升,精确类任务则普遍恶化,且影响持续存在。
- 适合研究模型推理机制或优化提示工程的研究者关注。
大型推理模型(LRMs)常提升数学与编程表现,但对指令遵循的影响尚不明确。本研究以Qwen3系列模型(1.7B-32B)为对象,在IFEval数据集上采用同权重的思维开启/关闭对照实验,四款Hunyuan模型提供跨家族方向性验证。总体通过率变化微小(-0.55至-3.52个百分点),但10%-20%的提示在模式切换中出现通过/失败反转,表明思考改变了错误模式——某些任务改善,某些恶化。后验分析显示,约束类型可划分为规划类(全局计数、结构、协调)与精确类(局部形式精准)。规划类在开启思考后整体性能提升,精确类则持续下降;该趋势在四个Hunyuan模型中方向一致。思考还影响最终答案长度,匹配长度分析虽显著缓解精确度下降,但仍存残余惩罚。通过交叉编码器相关性度量分析思考轨迹,发现:中立类呈现正相关(r≈0.15);规划类相关性接近零(r≈0.02),尽管轨迹参与度可观,反映轨迹相关性与最终合规性间存在执行鸿沟;精确类呈轻微负相关(r≈-0.05),失败样本的平均相关性高于通过样本。在四个模型规模(1.7B-14B)上进行激活修补实验表明,精确类翻转实例的恢复率高于规划类(均值32%-58% vs. 14%-40%),尤其在14B模型上差距达约30个百分点。
原文摘要 · Abstract (English)
Large reasoning models (LRMs) often improve math and coding performance, but their effect on instruction following is unclear. We study IFEval with Qwen3 models (1.7B-32B), using same-weights Thinking ON/OFF controls; four Hunyuan models provide directional cross-family support. Aggregate pass-rate changes are small (-0.55 to -3.52 pp), yet 10-20% of prompts switch between pass and fail across modes, suggesting that thinking changes the pattern of errors--some prompts improve while others worsen--rather than uniformly degrading performance. Under a post-hoc Qwen3-derived grouping, constraint types separate into Planning (global counting, structure, coordination), which improves at the class level under thinking, and Precision (exact local form), which consistently worsens; the class-level Planning/Precision sign pattern holds directionally for all four Hunyuan models despite Hunyuan's opposite aggregate direction. Thinking also changes final-answer length; matched-length analyses substantially reduce the Precision drop, but a residual penalty remains. Analyzing thinking traces with a cross-encoder relevance metric reveals three patterns: Neutral shows a positive relevance-compliance link (r approximately 0.15); Planning shows near-zero predictive correlation (r approximately 0.02) despite measurable trace engagement, consistent with an execution gap between CE-measured trace relevance and final-answer compliance; Precision shows a small negative correlation (r approximately -0.05), with failing instances having higher mean relevance than passing ones. Activation patching across four model sizes (1.7B-14B) shows that Precision flip instances are more often restored than Planning flip instances (32-58% vs. 14-40% mean layer-restoration), with the largest gap at 14B (about 30 pp).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。