arXiv:2503.05891cs.CL2025-03被引 4

用猜码游戏构建可扩展的推理评测,检验大模型真实逻辑能力

MastermindEval: A Simple But Scalable Reasoning Benchmark

  • 基于猜码游戏设计可解释的演绎推理测评
  • 简单题目仍难倒现有模型,且复杂度越高越易出错
  • 适合评估推理模型进化能力,尤其关注信息整合瓶颈

大语言模型在语言理解与数学任务上表现优异,但对其真实推理能力的评估仍面临挑战。为应对OpenAI o1、DeepSeek R1等先进推理模型的发展,本文提出MastermindEval——一个基于猜码游戏设计的简单、可扩展且可解释的演绎推理基准。该基准支持两种评估范式:(1)代理式评估,模型自主进行游戏;(2)演绎推理评估,给定已进行的游戏状态,模型需推断唯一正确密码。实验表明,即使简单的谜题对当前模型也极具挑战性,且随着需整合的信息量增加,模型推理能力显著下降。该基准具备未来扩展潜力,可追踪更先进模型的推理演进。

原文摘要 · Abstract (English)

Recent advancements in large language models (LLMs) have led to remarkable performance across a wide range of language understanding and mathematical tasks. As a result, increasing attention has been given to assessing the true reasoning capabilities of LLMs, driving research into commonsense, numerical, logical, and qualitative reasoning. However, with the rapid progress of reasoning-focused models such as OpenAI's o1 and DeepSeek's R1, there has been a growing demand for reasoning benchmarks that can keep pace with ongoing model developments. In this paper, we introduce MastermindEval, a simple, scalable, and interpretable deductive reasoning benchmark inspired by the board game Mastermind. Our benchmark supports two evaluation paradigms: (1) agentic evaluation, in which the model autonomously plays the game, and (2) deductive reasoning evaluation, in which the model is given a pre-played game state with only one possible valid code to infer. In our experimental results we (1) find that even easy Mastermind instances are difficult for current models and (2) demonstrate that the benchmark is scalable to possibly more advanced models in the future Furthermore, we investigate possible reasons why models cannot deduce the final solution and find that current models are limited in deducing the concealed code as the number of statement to combine information from is increasing.

推理评测逻辑推理大模型评估可扩展性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。