arXiv:2502.16906cs.CL2025-02被引 21

用程序验证生成开放题,更真实评估大模型逻辑推理能力

AutoLogi: Automated Generation of Logic Puzzles for Evaluating Reasoning Abilities of Large Language Models

  • 用程序自动构造可验证的开放式逻辑谜题
  • 模型得分从21%-37%扩大到35%-73%,区分度提升
  • 适合想真实测试或提升模型推理能力的研究者

尽管大语言模型的逻辑推理评估备受关注,现有基准多依赖易被随机猜测的多项选择题,导致性能高估和波动严重。为获得更准确的评估,我们提出一种自动生成开放式逻辑谜题的方法,并据此构建了双语基准AutoLogi。该方法结合程序化验证与可控难度设计,使评估更可靠,能更好区分模型真实推理能力。对八种主流LLM的评估显示,AutoLogi下模型得分范围为35%至73%,远超原多项选择数据集的21%至37%。此外,该合成方法可通过引入程序验证器进行拒绝采样,生成高质量训练数据,系统性提升模型在多种数据集上的推理能力。

原文摘要 · Abstract (English)

While logical reasoning evaluation of Large Language Models (LLMs) has attracted significant attention, existing benchmarks predominantly rely on multiple-choice formats that are vulnerable to random guessing, leading to overestimated performance and substantial performance fluctuations. To obtain more accurate assessments of models' reasoning capabilities, we propose an automated method for synthesizing open-ended logic puzzles, and use it to develop a bilingual benchmark, AutoLogi. Our approach features program-based verification and controllable difficulty levels, enabling more reliable evaluation that better distinguishes models' reasoning abilities. Extensive evaluation of eight modern LLMs shows that AutoLogi can better reflect true model capabilities, with performance scores spanning from 35% to 73% compared to the narrower range of 21% to 37% on the source multiple-choice dataset. Beyond benchmark creation, this synthesis method can generate high-quality training data by incorporating program verifiers into the rejection sampling process, enabling systematic enhancement of LLMs' reasoning capabilities across diverse datasets.

逻辑推理评测基准自动生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。