arXiv:2502.13125cs.CL2025-02被引 5

测试大模型识破逻辑谬误和误导前提的能力,发现表现普遍不佳。

RuozhiBench: Evaluating LLMs with Logical Fallacies and Misleading Premises

  • 构建了含677个问题的双语数据集,专攻误导性推理
  • 顶尖模型仅62%准确率,远低于人类超90%水平
  • 揭示主流模型在复杂推理中易受误导的弱点

近期大语言模型在复杂推理任务上表现突出,但其识别和应对包含逻辑谬误或故意误导前提文本的能力仍研究不足。为此,我们提出RuozhiBench,一个双语数据集,包含677个经人工精心设计并经专家评审的问题,涵盖多种欺骗性推理形式。我们在该数据集上对来自5个系列的17个大语言模型进行了全面评估,采用开放式与二选一两种格式,并深入分析评估协议与结果模式。尽管这些模型在传统基准上得分较高,但在识别和正确推理逻辑谬误方面能力有限,即使表现最佳的Claude-3-haiku模型也仅达62%准确率,而人类准确率超过90%。

原文摘要 · Abstract (English)

Recent advances in large language models (LLMs) have shown that they can answer questions requiring complex reasoning. However, their ability to identify and respond to text containing logical fallacies or deliberately misleading premises remains less studied. To address this gap, we introduce RuozhiBench, a bilingual dataset comprising 677 carefully curated questions that contain various forms of deceptive reasoning, meticulously crafted through extensive human effort and expert review. In a comprehensive evaluation of 17 LLMs from 5 Series over RuozhiBench using both open-ended and two-choice formats, we conduct extensive analyses on evaluation protocols and result patterns. Despite their high scores on conventional benchmarks, these models showed limited ability to detect and reason correctly about logical fallacies, with even the best-performing model, Claude-3-haiku, achieving only 62% accuracy compared to the human of more than 90%.

逻辑推理模型评测谬误检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。