用经典逻辑谜题测试AI组合推理能力
ClassicLogic: A Knowledge-Driven Benchmark of Classic Puzzle Games for Evaluating Compositional Generalization

- 构建四类谜题的分层知识库,明确策略组合关系
- 支持从基础规则到多步复合策略的渐进式评估
- 适合研究神经符号系统与复杂推理的学者
组合泛化——理解并生成已知组件新组合的能力——仍是现代人工智能的核心挑战。现有基准较少,且多集中于语言任务,缺乏复杂而明确的组合结构。我们提出ClassicLogic,一个用于评估智能体学习与组合求解策略能力的新基准套件。该套件包含四种经典逻辑谜题:数独(Sudoku)、肯肯(KenKen)、卡鲁科(Kakuro)和不等号填数(Futoshiki)。其核心创新在于为每种游戏构建分层、显式的知识库,将复杂求解策略形式化为更基础策略的组合。这一结构支持对智能体推理能力的细粒度评估,涵盖从学习基本规则到应用多步组合策略解决数学上验证难度递增谜题的全过程。开源基准为推进神经符号及其他先进AI推理系统提供了具有挑战性的新测试平台。
原文摘要 · Abstract (English)
Compositional generalization, the ability to understand and produce novel combinations of known components, remains a fundamental challenge for modern artificial intelligence. While few benchmarks exist, many focus on linguistic tasks and lack complex, explicit compositional structures. We introduce ClassicLogic, a new benchmark suite designed to evaluate an agent's ability to learn and compose problem-solving strategies. The benchmark consists of four classic logic puzzles: Sudoku, KenKen, Kakuro, and Futoshiki. Its core innovation is a hierarchical, explicit knowledge base for each game, where complex solving strategies are formally defined as compositions of simpler, foundational strategies. This structure allows for fine-grained evaluation of an agent's reasoning capabilities, from learning basic rules to applying multi-step compositional strategies to solve puzzles of increasing, mathematically validated difficulty. The open-source benchmark provides a challenging new testbed for advancing neuro-symbolic and other advanced AI reasoning systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。