arXiv:2608.12321cs.CLcs.AI2026-08

大模型知道规则却用不上,问题出在激活路径不通。

LLMs Know the Constraint But Do Not Use It: Activation Bottlenecks in Pragmatic Constraint Reasoning

  • 区分知识存储与决策激活,发现模型有约束知识但未调用
  • 14个模型测试显示,修复激活路径可提升6.4纳特性能
  • 适合研究模型推理机制与认知偏差的学者

当显著的表面线索与隐含的可行性约束冲突时,大语言模型常出现错误——但整体准确率混淆了真正的约束推理与保守默认行为。我们提出‘条件性约束激活’概念:约束被内部编码(知识),在有无约束提示下对称存在(对称性),但仅部分进入决策流程(路由),且可通过外部激活修复(修复性)。对14个模型的四联诊断揭示两种失败模式;两个开放权重探针解码约束成功率超88%,但激活修复仅在一个模式有效(+6.4纳特),另一个无效(-0.07)。在缓解前沿上,所有提示干预均未能突破瓶颈,反而通过‘前提提及’单一通路加剧保守偏见。隐藏约束失败本质是路由问题,而非知识缺失。

原文摘要 · Abstract (English)

When a salient surface cue competes with an implicit feasibility constraint, LLMs often fail -- but aggregate accuracy conflates genuine constraint inference with conservative defaulting. We formalize the distinction as conditional constraint activation: the constraint is internally encoded (Knowledge) symmetrically across constraint-present and -absent prompts (Symmetry), yet only sometimes routed into the decision (Routing) and repairable by a donor activation (Repair). A quartet diagnostic over 14 models reveals two failure modes; probes on two open weights decode the constraint above $88\%$, yet activation patching repairs one ($+6.4$ nats) and not the other ($-0.07$). On a mitigation frontier, no prompted intervention reaches the repair corner: all inflate conservative bias through a single mediation pathway -- prerequisite mention. Hidden-constraint failure is a routing problem, not a knowledge problem.

大模型推理约束推理激活路径

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。