中文逻辑推理新基准,用专家审核+Z3验证确保题目质量。
LLMEval-Logic: A Solver-Verified Chinese Benchmark for Logical Reasoning of LLMs with Adversarial Hardening

- 从真实场景构建题目,经专家审核与Z3验证答案正确性。
- 顶尖模型在难题集上准确率仅37.5%,形式化得分最高60.16%。
- 适合评估大模型逻辑推理能力,尤其关注鲁棒性和形式化能力。
评估大语言模型在自然语言逻辑推理任务中的表现至关重要,因为规则类任务要求结论严格遵循前提。现有逻辑推理基准多通过公式采样生成模板化语句,仅提供粗略或未经审计的形式化标注,且已被前沿模型迅速饱和。我们提出LLMEval-Logic,一个基于真实情境的中文逻辑推理基准。其流程包括:由领域作者生成自然语言题目并由专家审核其形式化标注,使用Z3验证答案正确性,构建自然语言到形式化的评分标准,并通过闭环对抗机制强化部分题目难度。基准以两组配对数据发布:246题的Base子集含1,400个专家制定的评分原子;190题的Hard子集包含938个封闭模型空间内的多步子问题。对14个前沿模型的评估显示显著差距:最佳模型在Hard子集上的准确率仅为37.5%,即使使用参考符号,模型在联合Z3+Rubric形式化评分中最高仅达60.16%。该基准已公开于https://github.com/llmeval/LLMEval-Logic。
原文摘要 · Abstract (English)
Evaluating large language models (LLMs) on natural-language logical reasoning is essential because rule-governed tasks require conclusions to follow strictly from stated premises. Many existing logical-reasoning benchmarks are generated by templating natural-language items from sampled formulas, provide only coarse or unaudited formal annotations, and are now quickly saturated by frontier reasoning models. We present LLMEval-Logic, a Chinese logical reasoning benchmark built from realistic situational scenarios. Its pipeline forward-authors and expert-audits natural-language items together with their reference formalizations, verifies annotated answers with Z3, constructs expert rubrics for natural-to-formal grading, and hardens selected items through a closed-loop adversarial workflow. The benchmark is released in two paired subsets: a 246-item Base subset shipped with 1,400 expert-developed rubric atoms, and a 190-item Hard subset with 938 multi-step sub-questions over closed model spaces. Evaluating 14 frontier LLMs on LLMEval-Logic reveals substantial gaps in current models: the best model reaches only 37.5% Hard Item Accuracy, and even with reference symbols the highest joint Z3+Rubric formalization score among evaluated models reaches only 60.16%. Our benchmark is publicly available at https://github.com/llmeval/LLMEval-Logic.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。