arXiv:2410.16502cs.CL2024-10ICML被引 4

测试大模型在逻辑与常识冲突时的人类式推理能力。

RULEBREAKERS: Challenging LLMs at the Crossroads between Formal Logic and Human-like Reasoning

  • 构建首个评估大模型识别逻辑悖论的基准数据集
  • 多数模型(含GPT-4o)准确率平庸,过度依赖规则
  • 揭示当前模型世界知识利用不足及注意力分布缺陷

形式逻辑通过符号化表达句子并应用规则推导结论,使计算机能以自然语言推理。然而,在本研究定义的“规则破坏者”场景中,该方法常产生人类基于常识和事实知识不会接受的结论。受认知科学启发,我们构建了RULEBREAKERS——首个严格评估大语言模型(LLMs)在人类式推理下识别和应对规则破坏者(而非非规则破坏者)能力的数据集。对七种大模型的评估显示,大多数模型(包括GPT-4o)在RULEBREAKERS上的准确率仅处于中等水平,且表现出过度僵化地应用逻辑规则的现象,与典型人类推理者预期不符。进一步分析表明,这种看似失败的表现可能与模型对世界知识的利用不足及其注意力分布模式有关。该研究揭示了当前大模型在推理上的局限性,也及时为近期依赖形式逻辑提升大模型通用推理能力的方法提供了警示,指出其可能加剧大模型与人类式推理之间的差距。

原文摘要 · Abstract (English)

Formal logic enables computers to reason in natural language by representing sentences in symbolic forms and applying rules to derive conclusions. However, in what our study characterizes as "rulebreaker" scenarios, this method can lead to conclusions that are typically not inferred or accepted by humans given their common sense and factual knowledge. Inspired by works in cognitive science, we create RULEBREAKERS, the first dataset for rigorously evaluating the ability of large language models (LLMs) to recognize and respond to rulebreakers (versus non-rulebreakers) in a human-like manner. Evaluating seven LLMs, we find that most models, including GPT-4o, achieve mediocre accuracy on RULEBREAKERS and exhibit some tendency to over-rigidly apply logical rules unlike what is expected from typical human reasoners. Further analysis suggests that this apparent failure is potentially associated with the models' poor utilization of their world knowledge and their attention distribution patterns. Whilst revealing a limitation of current LLMs, our study also provides a timely counterbalance to a growing body of recent works that propose methods relying on formal logic to improve LLMs' general reasoning capabilities, highlighting their risk of further increasing divergence between LLMs and human-like reasoning.

逻辑推理大模型评估常识推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。