arXiv:2603.02119cs.AIcs.GT2026-03被引 4

用铅笔谜题测试大模型多步可验证推理能力,精准定位错误。

Pencil Puzzle Bench: A Benchmark for Multi-Step Verifiable Reasoning

  • 基于6万多个谜题构建可逐步验证的评估框架
  • 模型在迭代验证下准确率最高提升至56.0%
  • 适合研究长程推理与过程监督的学者

我们提出Pencil Puzzle Bench,一个通过铅笔谜题评估大语言模型推理能力的框架。铅笔谜题是一类与NP完全问题密切相关的约束满足问题,具有确定性且支持逐步验证。从包含94种类型、共62,231个谜题的数据库中,我们筛选出涵盖20种类型的300个谜题作为基准测试集,评估了来自11家机构的51个模型,采用两种模式:直接提问(单次)和代理式(多轮迭代验证)。该基准的关键优势在于,每个中间棋盘状态均可根据特定类型规则进行校验,精确定位违反的具体规则,为过程监督和强化学习提供密集的每步奖励信号。评估发现模型能力存在两个维度:(1) 推理努力程度提升,GPT-5.2从无推理到最大努力时准确率提升81倍;(2) 代理式迭代能力,Claude Opus 4.6通过迭代检查使准确率从0.3%升至30.0%,GPT-5.2@xhigh则从20.2%升至56.0%。代理尝试中位数为29轮,耗时17分钟,最长超过1,221轮,持续14.3小时,是对长上下文利用能力的严峻考验。

原文摘要 · Abstract (English)

We introduce Pencil Puzzle Bench, a framework for evaluating large language model reasoning through pencil puzzles, a family of constraint-satisfaction problems closely related to NP-complete problems, with deterministic, step-level verification. From a database of 62,231 puzzles across 94 varieties with verified unique solutions, we select a benchmark of 300 puzzles spanning 20 varieties and evaluate 51 models from 11 providers in two modes: direct ask (single-shot) and agentic (multi-turn with iterative verification). A key differentiator of our benchmark is that every intermediate board state can be checked against variety-specific constraints, localizing errors to the exact rule violated, providing the infrastructure for dense, per-move reward signals for process supervision and reinforcement learning. Our evaluation reveals two distinct axes of capability: (1) reasoning effort scaling, where GPT-5.2 improves 81x from no reasoning to maximum effort; and (2) agentic iteration, where Claude Opus 4.6 rises from 0.3% to 30.0% through iterative checking, while GPT-5.2@xhigh improves from 20.2% to 56.0%. Agentic attempts span a median of 29 turns over 17 minutes, with the longest exceeding 1,221 turns and 14.3 hours - a demanding test of long-context utilization, not just reasoning.

推理评估多步推理过程监督谜题基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。