arXiv:2502.09589cs.CLcs.LO2025-02ACL被引 4

逻辑形式比概率更能解释大模型和人类的推理表现差异。

Logical forms complement probability in understanding language model (and human) performance

  • 构建命题与模态逻辑的假设及选言三段论测试集,控制变量研究推理能力。
  • 发现逻辑形式是预测大模型行为的关键因素,与输入概率同等重要。
  • 对比人类与大模型推理表现,揭示共性与差异,助力理解认知机制。

随着大语言模型(LLMs)在自然语言规划中的应用日益广泛,理解其行为变得尤为重要。本文系统研究了大模型在自然语言中的逻辑推理能力。我们引入了一个受控数据集,包含命题逻辑与模态逻辑的假设及选言三段论,并以此作为评估大模型性能的基准。结果表明,除了输入概率(Gonen et al., 2023;McCoy et al., 2024)外,逻辑形式同样是预测大模型行为的重要因素。此外,通过收集并比较人类与大模型的行为数据,我们揭示了两者在逻辑推理上的相似性与差异性,为理解语言模型与人类认知提供了新视角。

原文摘要 · Abstract (English)

With the increasing interest in using large language models (LLMs) for planning in natural language, understanding their behaviors becomes an important research question. This work conducts a systematic investigation of LLMs' ability to perform logical reasoning in natural language. We introduce a controlled dataset of hypothetical and disjunctive syllogisms in propositional and modal logic and use it as the testbed for understanding LLM performance. Our results lead to novel insights in predicting LLM behaviors: in addition to the probability of input (Gonen et al., 2023; McCoy et al., 2024), logical forms should be considered as important factors. In addition, we show similarities and discrepancies between the logical reasoning performances of humans and LLMs by collecting and comparing behavioral data from both.

逻辑推理大模型行为认知对比

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。