测试28个大模型对默认推理能力,发现表现参差不齐。
Generics and Default Reasoning in Large Language Models
- 用20种可反驳推理模式测试模型对通用语句的理解
- 多数模型在零样本下准确率超75%,但链式思考会下降11%以上
- 难区分默认推理与演绎推理,常把泛化语句当全称命题
本文评估了28个大型语言模型(LLMs)在20种涉及通用概括(如‘鸟会飞’、‘乌鸦是黑的’)的可反驳推理模式中的表现,这些模式是非单调逻辑的核心。通用语句因其允许例外的复杂行为,在语言学、哲学、逻辑学和认知科学中具有重要意义,是默认推理、认知和概念习得的关键。研究发现,尽管一些前沿模型在多数默认推理任务中表现良好,但不同模型及提示方式间表现差异显著。少样本提示对部分模型有小幅提升,但链式思考(CoT)提示常导致严重性能下降(在零样本准确率高于75%的模型中,平均下降11.14%,标准差15.74%,温度0)。大多数模型难以区分可反驳推理与演绎推理,或将通用语句误读为全称命题。这些发现揭示了当前大模型在默认推理方面的潜力与局限。
原文摘要 · Abstract (English)
This paper evaluates the capabilities of 28 large language models (LLMs) to reason with 20 defeasible reasoning patterns involving generic generalizations (e.g., 'Birds fly', 'Ravens are black') central to non-monotonic logic. Generics are of special interest to linguists, philosophers, logicians, and cognitive scientists because of their complex exception-permitting behaviour and their centrality to default reasoning, cognition, and concept acquisition. We find that while several frontier models handle many default reasoning problems well, performance varies widely across models and prompting styles. Few-shot prompting modestly improves performance for some models, but chain-of-thought (CoT) prompting often leads to serious performance degradation (mean accuracy drop -11.14%, SD 15.74% in models performing above 75% accuracy in zero-shot condition, temperature 0). Most models either struggle to distinguish between defeasible and deductive inference or misinterpret generics as universal statements. These findings underscore both the promise and limits of current LLMs for default reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。