用非常规数独题测试大模型的创造性多步逻辑推理能力
Sudoku-Bench: Evaluating creative reasoning with Sudoku variants

- 设计独特数独变体,避免记忆套路,逼模型找新解法
- 顶尖大模型独立解题成功率不足15%
- 适合研究长程逻辑推理与创意解题的学者
现有大语言模型推理评测基准常无法捕捉真实创造力,往往奖励对已有模式的记忆。为此,我们提出Sudoku-Bench,一个精心挑选的、具挑战性的数独变体集合,专门用于评估创造性、多步逻辑推理能力。数独变体具有独特或微妙交互的约束条件,使记忆不可行,要求求解者发现新的逻辑突破口(“break-ins”)。尽管形式多样,这些变体保持共同且紧凑的结构,便于清晰一致的评估。Sudoku-Bench包含精心筛选的题目集、标准化文本表示方式,以及兼容数千个公开可得题目的灵活工具,易于扩展为通用研究环境。基线实验显示,当前最先进的大模型在无辅助情况下解题成功率低于15%,凸显了提升长程战略推理能力的巨大空间。
原文摘要 · Abstract (English)
Existing reasoning benchmarks for large language models (LLMs) frequently fail to capture authentic creativity, often rewarding memorization of previously observed patterns. We address this shortcoming with Sudoku-Bench, a curated benchmark of challenging and unconventional Sudoku variants specifically selected to evaluate creative, multi-step logical reasoning. Sudoku variants form an unusually effective domain for reasoning research: each puzzle introduces unique or subtly interacting constraints, making memorization infeasible and requiring solvers to identify novel logical breakthroughs (``break-ins''). Despite their diversity, Sudoku variants maintain a common and compact structure, enabling clear and consistent evaluation. Sudoku-Bench includes a carefully chosen puzzle set, a standardized text-based puzzle representation, and flexible tools compatible with thousands of publicly available puzzles -- making it easy to extend into a general research environment. Baseline experiments show that state-of-the-art LLMs solve fewer than 15\% of puzzles unaided, highlighting significant opportunities to advance long-horizon, strategic reasoning capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。