arXiv:2411.12498cs.LGcs.AI2024-11NeurIPS被引 57

用自动生成的逻辑题训练大模型,显著提升推理能力。

Enhancing Reasoning Capabilities of LLMs via Principled Synthetic Logic Corpus

  • 基于符号逻辑设计高质量合成样本,构建多样化的推理数据集
  • 在逻辑、数学和编码任务上最高提升30分,整体推理能力显著增强
  • 适合想提升模型逻辑推理能力的研究者与开发者使用

大型语言模型虽能应对多种任务,但在推理方面仍表现不佳。为此,我们提出一种名为附加逻辑训练(ALT)的方法,通过程序生成的逻辑推理样本来增强模型的推理能力。首先,结合符号逻辑理论与已有实证经验,确立高质量样本的设计原则;随后,依据这些原则构建了一个名为形式化逻辑推理由多样化(FLD×2)的合成语料库,包含多步推导、未知事实、多样化推理规则、多样化语言表达及具有挑战性的干扰项。实验表明,在FLD×2上进行的ALT可显著提升当前先进大模型(如LLaMA-3.1-70B)的推理能力,逻辑推理基准测试最高提升30分,数学与编程基准最高提升10分,BBH基准套件提升5分。

原文摘要 · Abstract (English)

Large language models (LLMs) are capable of solving a wide range of tasks, yet they have struggled with reasoning. To address this, we propose $\textbf{Additional Logic Training (ALT)}$, which aims to enhance LLMs' reasoning capabilities by program-generated logical reasoning samples. We first establish principles for designing high-quality samples by integrating symbolic logic theory and previous empirical insights. Then, based on these principles, we construct a synthetic corpus named $\textbf{Formal Logic Deduction Diverse}$ ($\textbf{FLD}$$_{\times 2}$), comprising numerous samples of multi-step deduction with unknown facts, diverse reasoning rules, diverse linguistic expressions, and challenging distractors. Finally, we empirically show that ALT on FLD$_{\times2}$ substantially enhances the reasoning capabilities of state-of-the-art LLMs, including LLaMA-3.1-70B. Improvements include gains of up to 30 points on logical reasoning benchmarks, up to 10 points on math and coding benchmarks, and 5 points on the benchmark suite BBH.

大模型推理逻辑训练合成数据语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。