测试大模型能否像人一样理解条件句的隐含意义,发现仍不具备人类的语用推理能力。
Tracing the ongoing emergence of human-like reasoning in Large Language Models
- 对比25个大模型与人类在4种语言中对条件句的推理表现
- 多数模型能正确处理逻辑真值表,但忽略人类特有的语用推断
- 模型表现不受开源/闭源、架构类型影响,说明语用推理仍在发展中
人类能自然超越字面含义:例如‘如果你割草,我就给你五十美元’通常被理解为只有割草才给钱;而‘如果你饿了,炉子里有披萨’则暗示披萨存在,与是否饥饿无关。尽管大语言模型在多项任务上表现接近人类,但其是否具备人类式推理仍不明确。为此,我们开展了一项匹配人群的实验,评估25个大模型在四种语言中对条件句的推理表现,并与同等数量的人类对照。结果表明,人类通过语用推断丰富逻辑推理,而模型表现更不稳定:部分模型严格遵循条件句的真值表,但忽视语用推断;另一些则始终采用单一解释,符合规则但非人类式推理。总体而言,大模型是准确的语义操作者,却未能捕捉人类推理中的语用增强特征。更重要的是,模型的准确性既不能由是否开源、训练方向或架构类型预测,也无法提升,表明语用推理仍是人工智能认知工具包中正在涌现的能力。
原文摘要 · Abstract (English)
Humans effortlessly go beyond literal meanings: If you mow the lawn, I will give you fifty dollars, is typically understood as implying that the speaker will pay only if the lawn is mowed, whereas If you are hungry, there is pizza in the oven implies that pizza is available regardless of the hearers hunger. Large Language Models - LLMs - show human-like performance on many tasks, yet it remains unclear whether they reason like humans. To address this, we conducted a population-matching experiment assessing how twentyfive LLMs compute conditional inferences across four languages, compared to an equal number of humans per language. We find that humans enrich logical reasoning through pragmatic inferences across languages. Model behavior is more variable. Some LLMs perfectly follow the truth-table of conditionals but they ignore pragmatic inferences, while others deviate from the truth-table, adhering to a single interpretation across the board, thus reflecting accurate rule-based processing but not human-like reasoning. Overall, LLMs are accurate semantic operators, but fail to capture the pragmatic enrichments characteristic of human reasoning. Crucially, LLM accuracy is neither predicted nor boosted by open vs. closed status, training orientation, or architecture type, suggesting that pragmatic reasoning is still an emerging ability in the cognitive toolkit of artificial systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。