arXiv:2606.29278cs.AIcs.CL2026-06中稿 · the 1st Workshop o…

测试大模型在复杂推理任务中的极限表现,发现不同任务类型天花板差异巨大。

The Complexity Ceiling Benchmark: A Multi-Domain Evaluation of Sequential Reasoning Under Depth Scaling

  • 通过控制任务语义,仅改变推理步骤数,评估模型深度推理能力。
  • 三类任务中,最长仅能正确处理4.7步,其他任务可支持50步推理。
  • 模型错误答案中近15%源于中间推理错误,且参数量无法预测长期推理性能。

我们提出复杂度天花板基准(CCB),在可控条件下评估语言模型在推理步骤增加时的性能衰减。该基准固定任务语义内容,仅在三个结构不同的领域中变化推理深度N(5到50步):基于空间状态的追踪、抽象符号指针操作和传递性关系推理。在五种前沿及开源大模型上共完成6000次实验,发现存在一致的每步几何衰减模式,且各领域天花板显著不同:前两个领域最强模型在N=50时仍保持>0.92的成功率;第三类任务所有模型在N=5时即崩溃,最佳模型的50%成功率阈值仅为4.7步,尽管其初始成功率可达0.863。层级追踪指标(TFBC)显示,14.5%的正确答案来自错误的中间推理过程。强制详细状态记录未提升上限(McNemar检验p=1.000),且首次推理偏离的平均步骤数k*比参数量更能预测任务内准确率。CCB与几何衰减模型共同将模型长程推理表现简化为每个任务族一个可解释数值。

原文摘要 · Abstract (English)

We introduce the Complexity Ceiling Benchmark (CCB), a controlled evaluation of how language-model reasoning decays as the number of required sequential steps grows. CCB fixes the semantic content of a task and varies only its depth N in {5,...,50} across three structurally distinct regimes: grounded spatial state-tracking, abstract symbolic pointer manipulation, and transitive relational inference. Across 6,000 trials over five frontier and open-weight LLMs we find a consistent pattern of geometric per-step decay with widely separated domain ceilings: on the first two regimes the strongest models retain pd>0.92 across N=50; on the third every model collapses by N=5, with the best model's 50%-success horizon at H0.5~4.7 steps despite pd=0.863. A trace-level metric (TFBC) shows that 14.5% of correct answers across the benchmark are reached via incorrect intermediate reasoning. Forced verbose state-tracking does not move the ceiling (McNemar p=1.000), and the mean step at which reasoning first diverges, k*, predicts within-domain accuracy better than parameter count. CCB and the geometric decay model together reduce a model's long-horizon reasoning profile to one interpretable number per task family.

推理能力深度推理模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。