arXiv:2510.12563cs.AI2025-10被引 5

新基准挑战大模型逻辑推理,发现其依赖记忆而非真正理解。

HardcoreLogic: Challenging Large Reasoning Models with Long-tail Logic Puzzle Games

  • 通过三类改造构建5000+罕见逻辑谜题,测试模型泛化能力。
  • 主流模型在新基准上性能大幅下降,暴露对记忆模式的依赖。
  • 适合研究模型真实推理能力与长尾泛化问题的学者。

大型推理模型(LRMs)在复杂任务中表现优异,包括需满足所有约束条件的逻辑谜题。然而,它们能否灵活应对非标准游戏变体仍不明确。现有数据集集中于9x9数独等常见谜题,易导致模型过拟合标准格式并记忆解法模式,掩盖了对新规则的理解缺陷。为此,我们提出HardcoreLogic,一个包含10种游戏、超过5000个谜题的挑战性基准,旨在测试LRMs在逻辑谜题‘长尾’上的鲁棒性。该基准通过三个维度系统改造经典谜题:复杂度提升(IC)、非常规元素(UE)和无解谜题(UP),降低对捷径记忆的依赖。对多种LRMs的评估显示,即使在现有基准上表现优异的模型,性能也显著下降,表明其严重依赖记忆化的刻板印象。复杂度提升是主要难度来源,但模型对细微规则变化亦表现不佳,即使未增加谜题难度。对可解与不可解谜题的系统误差分析进一步揭示了真实推理能力的不足。总体而言,HardcoreLogic暴露出当前LRMs的局限性,并为推进高级逻辑推理建立了新基准。

原文摘要 · Abstract (English)

Large Reasoning Models (LRMs) have demonstrated impressive performance on complex tasks, including logical puzzle games that require deriving solutions satisfying all constraints. However, whether they can flexibly apply appropriate rules to varying conditions, particularly when faced with non-canonical game variants, remains an open question. Existing corpora focus on popular puzzles like 9x9 Sudoku, risking overfitting to canonical formats and memorization of solution patterns, which can mask deficiencies in understanding novel rules or adapting strategies to new variants. To address this, we introduce HardcoreLogic, a challenging benchmark of over 5,000 puzzles across 10 games, designed to test the robustness of LRMs on the "long-tail" of logical games. HardcoreLogic systematically transforms canonical puzzles through three dimensions: Increased Complexity (IC), Uncommon Elements (UE), and Unsolvable Puzzles (UP), reducing reliance on shortcut memorization. Evaluations on a diverse set of LRMs reveal significant performance drops, even for models achieving top scores on existing benchmarks, indicating heavy reliance on memorized stereotypes. While increased complexity is the dominant source of difficulty, models also struggle with subtle rule variations that do not necessarily increase puzzle difficulty. Our systematic error analysis on solvable and unsolvable puzzles further highlights gaps in genuine reasoning. Overall, HardcoreLogic exposes the limitations of current LRMs and establishes a benchmark for advancing high-level logical reasoning.

逻辑推理大模型评估长尾泛化基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。