大模型常被表面线索误导,忽略隐含约束,本文揭示其推理漏洞并提出干预方案。
The Model Says Walk: How Surface Heuristics Override Implicit Constraints in LLM Reasoning
- 构建覆盖4类启发式与5类约束的评测基准HOB,设计最小差异对与显性梯度
- 14个模型平均仅44%正确处理存在约束,提示缺失导致性能下降15个百分点
- 明确分解目标可提升5.0个百分点,证明约束枚举是关键干预手段
大型语言模型在显著表面线索与隐含可行性约束冲突时会失效。本文提出启发式覆盖基准(HOB):包含500个实例,覆盖4类启发式与5类约束,具备最小差异对和显性梯度。结合诊断-测量-桥接-治疗的可验证行为分析框架,对六种模型在洗车问题上的因果行为分析显示,距离线索的影响是目标的8.7至38倍,归因更符合关键词关联而非组合推理。在14个模型中严格10/10评估下,无一模型超过75%,存在约束最难处理,正确率仅44%。仅添加最小提示即可提升15个百分点,表明是推理缺陷而非知识缺失。然而,移除约束后12个模型表现下降最多达39个百分点,暴露保守偏差。对Gemini 3.1 Pro的思维模式消融实验显示,开启思考从74.6%降至58.4%,而显式目标分解可恢复至71.2%。因此,内部反思有实际作用,显式提示可部分替代。推理模式不显著优于非推理模型:控制能力等级后,推理残差效应仅1.8个百分点且不显著。参数探针显示,该符号函数模式也适用于成本、效率与语义相似性启发式。目标分解提示比通用链式思考提升5.0个百分点,优于后者3.1个百分点,确认约束枚举是核心机制。总体而言,启发式覆盖是系统性推理漏洞,其根源在于推理顺序而非知识,且已有可验证干预方案。
原文摘要 · Abstract (English)
Large language models fail when a salient surface cue conflicts with an unstated feasibility constraint. We introduce the Heuristic Override Benchmark (HOB): 500 instances spanning 4 heuristic families and 5 constraint families, with minimal pairs and explicitness gradients. We pair HOB with a falsifiable behavioral characterization following a diagnose-measure-bridge-treat arc. Causal-behavioral analysis of the car wash problem across six models reveals context-independent sigmoid heuristics: the distance cue has 8.7 to 38 times more influence than the goal, and attribution better matches keyword association than compositional inference. Across 14 models, strict 10/10 evaluation shows that no model exceeds 75%, and presence constraints are hardest at 44%. A minimal hint improves performance by 15 pp, suggesting a constraint-inference failure rather than missing knowledge. However, 12 of 14 models perform worse when the constraint is removed, by up to 39 pp, revealing conservative bias. A thinking-mode ablation on Gemini 3.1 Pro drops performance from 74.6% with thinking on to 58.4% with thinking off, while explicit goal decomposition recovers it to 71.2%. Thus, internal deliberation does useful work, and explicit prompting can partially substitute for it. Reasoning models do not categorically outperform non-reasoning peers: after controlling for capability rank, the residual reasoning-mode effect is 1.8 pp and is not significant. Parametric probes show that the sigmoid pattern generalizes to cost, efficiency, and semantic-similarity heuristics. Goal-decomposition prompting improves performance by 5.0 pp, compared with 3.1 pp for generic chain-of-thought, isolating constraint enumeration as the active ingredient. Overall, heuristic override is a systematic reasoning vulnerability with a quantified locus in inference order, not knowledge, and a tested intervention.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。