提出用复杂度外分布评估推理能力,打破模型伪装的智能幻觉。
Bridging Reasoning to Learning: Unmasking Illusions using Complexity Out of Distribution Generalization
- 以解题复杂度(如步骤数、结构丰富度)定义推理泛化能力
- 模型需在训练未见的高复杂度任务上保持性能才算真正具备推理能力
- 为构建可信赖的推理模型提供方法论,适合研究通用AI的学者
当前人工智能正从模式识别迈向需要逐步推理的系统2型任务,尤其在大语言模型中表现突出。然而,与学习中的泛化和分布外(OoD)评估已有明确定义不同,推理能力尚无清晰统一的衡量标准。本文提出复杂度外分布(Complexity OoD)泛化框架,用于定义与测量推理能力:当模型在测试实例的最小解题复杂度(包括表示层面的结构复杂性或计算层面的推理步数/程序长度)超过所有训练样本时仍能保持性能,则具备复杂度OoD泛化能力。通过解题描述的科尔莫戈罗夫复杂度及其操作代理(如对象/关系数量、推理步骤数)进行形式化,明确区分了复杂度OoD与长度和组合式OoD。该视角统一了学习与推理:低复杂度下可系统1完成的任务,在复杂度压力下转为系统2型,而系统2可视为对解结构的泛化。本文进一步给出实践建议:将复杂度纳入基准与评估指标设计,转向追踪解题轨迹的监督方式,探索支持复杂度泛化的归纳偏置,应对推理学习中的捷径依赖、语义鲁棒性、灾难性遗忘与步骤校准等问题。由于复杂度OoD无法仅靠数据量扩展解决,实现稳健推理需架构与训练策略显式建模并分配计算资源以适应复杂度变化。
原文摘要 · Abstract (English)
Recent progress has pushed AI frontiers from pattern recognition tasks toward problems that require step by step, System2 style reasoning, especially with large language models. Yet, unlike learning, where generalization and out of distribution (OoD) evaluation concepts are well formalized, there is no clear, consistent definition or metric for reasoning ability. We propose Complexity Out of Distribution (Complexity OoD) generalization as a framework and problem setting to define and measure reasoning. A model exhibits Complexity OoD generalization when it maintains performance on test instances whose minimal required solution complexity, either representational (richer solution structure) or computational (more reasoning steps/program length), exceeds that of all training examples. We formalize complexity via solution description Kolmogorov complexity and operational proxies (e.g., object/relation counts; reasoning step counts), clarifying how Complexity OoD differs from length and compositional OoD. This lens unifies learning and reasoning: many cases solvable with System1 like processing at low complexity become System2 like under complexity pressure, while System2 can be viewed as generalization over solution structures. We translate this perspective into practice with recommendations for operationalizing Complexity OoD across the stack: incorporating complexity into benchmark and evaluation metric design, rethinking supervision to target solution traces, seeking and designing inductive biases for Complexity OoD generalization, addressing learning to reason spillovers such as spurious shortcuts, semantic robustness, catastrophic forgetting, and step wise calibration. Because Complexity OoD cannot be solved by scaling data alone, progress toward robust reasoning will require architectures and training regimes that explicitly model and allocate computation with respect to complexity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。