arXiv:2503.23487cs.AIcs.CL2025-03ACL被引 8

LLM推理模型在多路径推理上表现差,本质是浅层析取推理。

Large Language and Reasoning Models are Shallow Disjunctive Reasoners

  • 用系统性关系组合测试推理能力,控制难度评估泛化性能
  • 单路径任务中推理模型优于普通LLM,但多路径任务完全崩溃
  • 发现其推理机制为浅层析取,无法真正处理复杂逻辑组合

大型语言模型(LLMs)在系统性推理上表现不佳,即使在看似表现良好的任务中,其成功也依赖于捷径而非真实推理能力,导致在分布外(OOD)样本上迅速失效。基于强化学习和思维链提示的后训练策略近期被视为重大突破。然而,这些方法在数学与编程之外的任务中的潜力仍不明确,尤其是缺乏真正的OOD问题。本文聚焦需要系统性关系组合的定性空间与时间推理任务,可精细控制问题难度以精确测量OOD泛化能力。结果发现,零样本推理模型(LRMs)在单路径推理任务中普遍优于对应LLM,但在多路径设置下表现不佳。微调后的LLM同样无法实现多路径泛化。我们还提供了行为证据表明,这类模型本质上是浅层析取推理者。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have been found to struggle with systematic reasoning. Even on tasks where they appear to perform well, their performance often depends on shortcuts, rather than on genuine reasoning abilities, leading them to collapse on out-of-distribution (OOD) examples. Post-training strategies based on reinforcement learning and chain-of-thought prompting have recently been hailed as a step change. However, little is known about the potential of the resulting ``Large Reasoning Models'' (LRMs) beyond maths and programming-based problem solving, where genuine OOD problems can be sparse. In this paper, we focus on tasks that require systematic relational composition for qualitative spatial and temporal reasoning. The setting allows fine control over problem difficulty to precisely measure OOD generalization. We find that, zero-shot LRMs generally outperform their LLM counterparts in single-path reasoning tasks but struggle in the multi-path setting. Whilst showing comparatively better results, fine-tuned LLMs are also not capable of multi-path generalization. We also provide evidence for the behavioral interpretation for this, i.e., that LRMs are shallow disjunctive reasoners.

推理模型多路径推理泛化能力浅层推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。