用逻辑编程生成真实语境下的谬误测试集,评估大模型的推理能力。
Socrates or Smartypants: Testing Logic Reasoning Capabilities of Large Language Models with Logic Programming-based Test Oracles
- 用Prolog规则自动生成含逻辑谬误的自然语言句子。
- 生成内容在微妙程度和质量上接近人工标注,且优于基线方法。
- 揭示过多推理步骤降低谬误识别准确率,结构化推理提升分类表现。
大型语言模型(LLMs)在语言理解与推理方面取得了显著进展,因此评估其逻辑推理能力变得至关重要。然而,现有数据集和基准测试往往局限于过于简单、不自然或情境受限的例子。为应对这一需求,我们提出了SmartyPat-Bench,一个基于真实高质量Reddit帖子中微妙逻辑谬误构建的挑战性、自然表达且系统标注的基准测试。该数据集提供更详细的谬误标注并具有更丰富的多样性。为克服人工数据收集与标注的局限性(如谬误类型失衡、劳动密集),我们引入SmartyPat——一种基于逻辑编程推理器的自动化框架。该框架利用Prolog规则系统生成逻辑谬误语句,再通过LLM将其转化为流畅的自然语言,确保谬误表征精确。大量实验表明,SmartyPat生成的谬误在微妙性和质量上可媲美人工内容,并显著优于基线方法。最终实验揭示了关于LLM能力的细致洞察:过多推理步骤会降低谬误检测准确率,而结构化推理则能提升谬误分类性能。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have achieved significant progress in language understanding and reasoning. Evaluating and analyzing their logical reasoning abilities has therefore become essential. However, existing datasets and benchmarks are often limited to overly simplistic, unnatural, or contextually constrained examples. In response to the growing demand, we introduce SmartyPat-Bench, a challenging, naturally expressed, and systematically labeled benchmark derived from real-world high-quality Reddit posts containing subtle logical fallacies. Unlike existing datasets and benchmarks, it provides more detailed annotations of logical fallacies and features more diverse data. To further scale up the study and address the limitations of manual data collection and labeling - such as fallacy-type imbalance and labor-intensive annotation - we introduce SmartyPat, an automated framework powered by logic programming-based oracles. SmartyPat utilizes Prolog rules to systematically generate logically fallacious statements, which are then refined into fluent natural-language sentences by LLMs, ensuring precise fallacy representation. Extensive evaluation demonstrates that SmartyPat produces fallacies comparable in subtlety and quality to human-generated content and significantly outperforms baseline methods. Finally, experiments reveal nuanced insights into LLM capabilities, highlighting that while excessive reasoning steps hinder fallacy detection accuracy, structured reasoning enhances fallacy categorization performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。