测试大模型是否像人类一样偏好简单解释,发现它们在复杂场景下表现不佳。
Do Language Models Follow Occam's Razor? An Evaluation of Parsimony in Inductive and Abductive Reasoning
- 设计新框架生成需归纳与溯因结合的推理题,支持一阶逻辑表达
- 发现顶尖模型在复杂世界中难以生成既正确又简洁的假设
- 提出量化简洁性的自动评估指标,适合研究模型思维偏好
非演绎推理(包括归纳与溯因)是解决复杂现实问题的关键。这类推理常存在多个有效假设,最简单的(符合奥卡姆剃刀原则)通常最有用。然而,现有大语言模型(LLMs)评估研究忽略了这一特性。本文填补该空白,探究LLMs的归纳与溯因能力是否遵循奥卡姆剃刀,并考察其推理正确性。为此,我们提出一个合成生成推理题的框架,能同时要求归纳与溯因推理,且可扩展至所有一阶逻辑表达的推理问题;任务为在给定世界模型下,生成解释观察结果的假设。我们还引入一种新自动化度量,评估假设在简洁性上的符合程度——既正确又最简的为高质量。实验结果表明,当前最优LLMs在简单场景下能完成归纳与溯因推理,但在复杂世界模型下难以生成高质量假设,即使采用提示学习和RLVR等增强技术亦然。
原文摘要 · Abstract (English)
Non-deductive reasoning, encompassing inductive and abductive reasoning, is essential in addressing complex real-world questions. One key feature of inductive and abductive reasoning is that there are many valid hypotheses; the simplest ones (those that adhere to Occam's Razor) are often most useful. However, this aspect is ignored in recent work that evaluates the non-deductive reasoning capabilities of large language models (LLMs). This work fills this gap, focusing on understanding whether the inductive and abductive reasoning capabilities of LLMs adhere to Occam's Razor, while also examining the correctness of their reasoning. To accomplish this goal, we introduce a framework to synthetically generate reasoning questions that (a) require inductive reasoning and abductive reasoning simultaneously; (b) is readily extended to produce any abductive/inductive reasoning question expressible in first-order logic. The task for the intelligent agent is to produce hypotheses to explain observations under a given world model. We also propose a new automated metric to assess whether hypotheses quantitatively adhere to Occam's Razor; those hypotheses that are correct and simplest are considered high-quality. Our findings on state-of-the-art LLMs suggest that LLMs can perform inductive and abductive reasoning in simple scenarios, but struggle with complex world models and with producing high-quality hypotheses, even with popular reasoning-enhancing techniques such as in-context learning and RLVR.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。