arXiv:2509.15587cs.CLcs.AI2025-09EMNLP被引 6

构建新基准评估大模型逻辑推理能力,避免多技能混淆与语言偏差。

DivLogicEval: A Framework for Benchmarking Logical Reasoning Evaluation in Large Language Models

  • 设计反直觉自然语句的逻辑题库,分离单一推理能力测试。
  • 发现大模型在复杂逻辑任务中表现显著下降,最优模型仅达62%准确率。
  • 适合研究模型推理缺陷或开发鲁棒性评估方法的学者使用。

自然语言中的逻辑推理被视为衡量大型语言模型(LLMs)智能水平的重要指标。现有主流基准常混合多种推理技能,导致对逻辑推理能力的评估失真;同时,现有逻辑推理基准在语言多样性上不足,其数据分布偏离理想逻辑推理评测标准,可能引发评估偏差。为此,本文提出全新经典逻辑推理基准DivLogicEval,由多样化的反直觉自然语句构成。为提升评估可靠性,还引入新评价指标,以缓解大模型固有的偏差与随机性影响。通过实验验证了DivLogicEval中问题所需的逻辑推理强度,并对比了不同主流LLMs在逻辑推理上的表现。

原文摘要 · Abstract (English)

Logic reasoning in natural language has been recognized as an important measure of human intelligence for Large Language Models (LLMs). Popular benchmarks may entangle multiple reasoning skills and thus provide unfaithful evaluations on the logic reasoning skill. Meanwhile, existing logic reasoning benchmarks are limited in language diversity and their distributions are deviated from the distribution of an ideal logic reasoning benchmark, which may lead to biased evaluation results. This paper thereby proposes a new classical logic benchmark DivLogicEval, consisting of natural sentences composed of diverse statements in a counterintuitive way. To ensure a more reliable evaluation, we also introduce a new evaluation metric that mitigates the influence of bias and randomness inherent in LLMs. Through experiments, we demonstrate the extent to which logical reasoning is required to answer the questions in DivLogicEval and compare the performance of different popular LLMs in conducting logical reasoning.

逻辑推理模型评估基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。