首次为大模型的'思维跃迁'提供形式化定义与测量方法。
When the Canonical Completion Is Wrong: Formalizing and Measuring the Jump in Large Language Models

- 用范畴论定义'跃迁':需突破默认推理路径并给出可验证的新解
- 9个认证实例测试中,4个前沿模型在248次实验中零次返回默认答案
- 跃迁失败源于推理资源耗尽或约束错误,非不愿放弃旧范式
大语言模型能否从证据跃迁至新公理体系(即'跳变')引发争议。本文提出四步形式化框架,定义跃迁为具有机器可验证证书的有限扩展问题:存在唯一正确解,且不同于由左/右凯恩延拓给出的默认完成。该默认解即模型无约束时的输出。证明跃迁实例良态且可生成无限难度案例。在9个认证实例、4个前沿模型上测试,248次受限实验中,凯恩默认率均为0,表明模型始终能主动放弃默认解;更高难度失败源于推理预算耗尽或约束错误,而非回退到默认。结果说明第二步跃迁非瓶颈。若能力缺失真实存在,根源可能在于构建约束或创设框架。
原文摘要 · Abstract (English)
Whether large language models (LLMs) can perform the abductive leap from evidence to a new system of axioms, commonly referred to as a jump, has recently attracted considerable debate. A prominent position holds that LLMs are structurally incapable of such jumps, while recent studies challenge both its mechanism and its evidence. However, the debate remains difficult to settle, since the field still lacks a formal definition of the jump and a measure to test either side. In this paper, we develop a formal account of the jump in four steps and measure the second. The steps ask what the default completion of partial data is, when abandoning it is forced, when the abandonment is correct, and how successive jumps compound. Specifically, we define a jump instance as a finite extension problem with a machine-checked certificate that a correct completion exists, is unique up to renaming, and differs from the canonical completion of the data. The canonical completion is given by the left and right Kan extensions and is also what models produce without constraints, so it serves as the default. We prove that jump instances are well-posed and establish a family theorem that certifies instances of unbounded difficulty without enumeration. We further formalize when a jump is correct and how successive jumps compound. Finally, we run the measurement on nine certified instances and four frontier models. The Kan-default rate is zero in all 248 constrained trials, so the models do jump at this step and abandon the excluded default every time. Failures at higher difficulty stem from exhausted reasoning budgets or constraint errors, never from reverting to the default. These results indicate that the second step is not the bottleneck. If the disputed incapacity is real, it lies in generating the constraints or inventing the framework. Code can be found at: https://github.com/EEthanShi/kan-jump-test.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。