对比大模型的直觉与逻辑推理,发现不同场景下各有优劣。
LogiDynamics: Unraveling the Dynamics of Inductive, Abductive and Deductive Logical Inferences in LLM Reasoning
- 设计可控实验,比较直觉(System 1)与逻辑(System 2)推理
- 视觉/符号任务中系统2表现更优,文本简单任务中系统1可胜出
- 自由作答任务中系统1反而更适应规则执行,适合优化策略部署
现代大语言模型采用多种逻辑推理机制,其策略性优化对提升能力至关重要。本文系统研究了归纳(系统1)与溯因/演绎(系统2)推理在大模型中的动态差异。通过控制类比推理环境,改变模态(文本、视觉、符号)、难度及任务格式(选择题/自由作答),发现系统2在视觉/符号模态和高难度任务中普遍更优,而系统1在文本和简单任务中具有竞争力。关键的是,任务格式显著影响二者相对优势,系统1在自由作答规则执行任务中甚至优于系统2。这些发现可推广至更广泛的上下文学习。此外,我们证明先进系统2策略如假设选择与迭代优化能显著提升大模型推理能力。本研究提供基础洞见与可操作指南,助力战略性部署逻辑推理以增强大模型表现。资源已开源:https://github.com/HKUST-KnowComp/LogiDynamics。
原文摘要 · Abstract (English)
Modern large language models (LLMs) employ diverse logical inference mechanisms for reasoning, making the strategic optimization of these approaches critical for advancing their capabilities. This paper systematically investigate the comparative dynamics of inductive (System 1) versus abductive/deductive (System 2) inference in LLMs. We utilize a controlled analogical reasoning environment, varying modality (textual, visual, symbolic), difficulty, and task format (MCQ / free-text). Our analysis reveals System 2 pipelines generally excel, particularly in visual/symbolic modalities and harder tasks, while System 1 is competitive for textual and easier problems. Crucially, task format significantly influences their relative advantage, with System 1 sometimes outperforming System 2 in free-text rule-execution. These core findings generalize to broader in-context learning. Furthermore, we demonstrate that advanced System 2 strategies like hypothesis selection and iterative refinement can substantially scale LLM reasoning. This study offers foundational insights and actionable guidelines for strategically deploying logical inference to enhance LLM reasoning. Resources are available at https://github.com/HKUST-KnowComp/LogiDynamics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。