arXiv:2410.09528cs.LGcs.AI2024-10被引 4

用自动构造的推理数据提升大模型的多步逻辑推理能力

Boosting Deductive Reasoning with Step Signals In RLHF

  • 基于形式逻辑生成可控复杂度的多步推理数据
  • 经强化学习微调后,模型在域内与域外任务上推理能力显著提升
  • 适合研究大模型推理能力或需要高质量推理训练数据的团队

逻辑推理是大型语言模型应对复杂问题的关键能力。多步推理任务尤其具有挑战性。基于形式逻辑理论,我们提出了一种自动化方法 Multi-step Deduction (MuseD),用于生成演绎推理数据。该方法可生成用于训练和测试的多步推理数据集,并能控制指令复杂度,支持不同难度下的模型训练与评估。通过强化学习人类反馈(RLHF)训练,所生成的数据使模型在域内与域外推理任务中均表现出显著提升。此外,我们还对多种模型的多步推理能力进行了测试。

原文摘要 · Abstract (English)

Logical reasoning is a crucial task for Large Language Models (LLMs), enabling them to tackle complex problems. Among reasoning tasks, multi-step reasoning poses a particular challenge. Grounded in the theory of formal logic, we have developed an automated method, Multi-step Deduction (MuseD), for deductive reasoning data. MuseD has allowed us to create training and testing datasets for multi-step reasoning. Our generation method enables control over the complexity of the generated instructions, facilitating training and evaluation of models across different difficulty levels. Through RLHF training, our training data has demonstrated significant improvements in logical capabilities for both in-domain of out-of-domain reasoning tasks. Additionally, we have conducted tests to assess the multi-step reasoning abilities of various models.

逻辑推理RLHF数据生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。