用三段论测试大模型推理能力,发现其组合性弱于递归性,提出神经符号混合模型解决。
Hybrid Models for Natural Language Reasoning: The Case of Syllogistic Logic
- 以三段论为基准,区分组合性与递归性两种泛化能力。
- 大模型在递归推理上表现良好,但组合性能力参差不齐,准确率从接近完美到显著偏低。
- 提出神经符号混合架构,兼顾效率与推理完备性,适合逻辑推理任务研究者。
尽管神经模型进展显著,其泛化能力——如逻辑推理应用中的核心需求——仍是关键挑战。我们明确泛化的两个基本方面:组合性(抽象复杂推理中基础逻辑规则的能力)和递归性(通过迭代应用推理规则构建复杂表示的能力)。文献中常将二者混同于泛化概念。为澄清此区别,我们以三段论片段为自然语言推理基准,评估大语言模型(LLMs)的逻辑泛化能力。我们扩展经典三段论形式,构建更复杂的结构,形成一个基础且表达力强的形式逻辑子集,支持对核心推理能力的可控评估。实验结果表明,在这一非平凡基准上,虽然大模型在递归性方面表现出合理能力,但在组合性上存在明显不足。这种差异并非均匀分布,详细分析显示不同三段论类型间的泛化性能差异显著,准确率从近乎完美到显著偏低不等。为克服这些局限并建立可靠的逻辑证明器,我们提出一种融合符号推理与神经计算的混合架构。该协同机制实现鲁棒高效的推理:神经组件加速处理,符号推理保障完备性。实验进一步表明,即使使用相对较小的神经组件,仍能保持高效率。总体而言,我们的分析既为神经符号混合方法提供了依据,也证明其在突破神经推理系统关键泛化障碍方面的潜力。
原文摘要 · Abstract (English)
Despite the remarkable progress in neural models, their ability to generalize, a cornerstone for applications such as logical reasoning, remains a critical challenge. We delineate two fundamental aspects of this ability: compositionality, the capacity to abstract atomic logical rules underlying complex inferences, and recursiveness, the aptitude to build intricate representations through iterative application of inference rules. In the literature, these two aspects are often conflated under the umbrella term of generalization. To sharpen this distinction, we investigate the logical generalization capabilities of LLMs using the syllogistic fragment as a benchmark for natural language reasoning. We extend classical syllogistic forms to construct more complex structures, yielding a foundational yet expressive subset of formal logic that supports controlled evaluation of essential reasoning abilities. Our findings on this non-trivial benchmark show that, while LLMs demonstrate reasonable proficiency in recursiveness, they struggle with compositionality. This disparity is not uniform, as a more detailed analysis reveals substantial variability in generalization performance across individual syllogistic types, ranging from near-perfect accuracy to significantly lower performance. To overcome these limitations and establish a reliable logical prover, we propose a hybrid architecture integrating symbolic reasoning with neural computation. This synergistic interaction enables robust and efficient inference, neural components accelerate processing, while symbolic reasoning guarantees completeness. Our experiments further show that high efficiency is preserved even when using relatively small neural components. Overall, our analysis provides both a rationale for hybrid neuro-symbolic approaches and evidence of their potential to address key generalization barriers in neural reasoning systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。