模型总按老套路走,即使用户给新指令也不听。
Reasoning Model is Stubborn: Diagnosing Instruction Overriding in Reasoning Models
- 构建专用诊断集,测试模型能否摆脱习惯性推理
- 发现三种错误模式:误解指令、不信输入、只听部分要求
- 适合研究大模型推理偏差与可控性的人看
大型语言模型在复杂推理任务中表现优异,但常依赖熟悉推理模式,即所谓‘推理僵化’。即使用户明确给出新条件,模型仍会默认使用惯用路径,导致错误结论。该问题在数学和逻辑谜题等需严格遵守约束的领域尤为严重。为系统研究这一现象,我们构建了专家设计的诊断数据集 dataset{},包含对现有数学基准 AIME 和 MATH500 的特殊改写版本,以及刻意重构的知名谜题,强制模型偏离原有策略。通过该数据集,我们识别出三类典型干扰模式:(i) 解释过载,(ii) 输入不信任,(iii) 部分指令关注,均导致模型忽略或扭曲用户指令。我们已公开该诊断集,以推动缓解语言模型推理僵化的研究。
原文摘要 · Abstract (English)
Large language models have demonstrated remarkable proficiency in long and complex reasoning tasks. However, they frequently exhibit a problematic reliance on familiar reasoning patterns, a phenomenon we term \textit{reasoning rigidity}. Despite explicit instructions from users, these models often override clearly stated conditions and default to habitual reasoning trajectories, leading to incorrect conclusions. This behavior presents significant challenges, particularly in domains such as mathematics and logic puzzle, where precise adherence to specified constraints is critical. To systematically investigate reasoning rigidity, a behavior largely unexplored in prior work, we introduce a expert-curated diagnostic set, \dataset{}. Our dataset includes specially modified variants of existing mathematical benchmarks, namely AIME and MATH500, as well as well-known puzzles deliberately redesigned to require deviation from familiar reasoning strategies. Using this dataset, we identify recurring contamination patterns that occur when models default to ingrained reasoning. Specifically, we categorize this contamination into three distinctive modes: (i) Interpretation Overload, (ii) Input Distrust, and (iii) Partial Instruction Attention, each causing models to ignore or distort provided instructions. We publicly release our diagnostic set to facilitate future research on mitigating reasoning rigidity in language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。