让大模型学会补全隐含推理步骤,提升逻辑准确性。
RATIONALYST: Mining Implicit Rationales for Process Supervision of Reasoning
- 从海量无标注数据中自动提取79,000条隐含理由,用于预训练
- 在7个推理基准上平均提升3.9%准确率,优于大模型验证器
- 适合需要可靠推理过程的AI系统开发者使用
大语言模型生成的推理步骤可能不完整,因其预训练数据中常省略隐含理由。为此,我们提出RATIONALYST,基于从大规模无标注数据(包括The Pile和多个推理数据集)中提取的79,000条理由进行预训练。该模型在多样化的推理任务中表现稳定,涵盖数学、常识、科学与逻辑推理。以LLaMa-3-8B为基础微调后,其在7个代表性推理基准上的平均准确率提升3.9%。相较于更大型的验证模型如GPT-4,以及同规模微调模型,RATIONALYST表现出更优性能。
原文摘要 · Abstract (English)
The reasoning steps generated by LLMs might be incomplete, as they mimic logical leaps common in everyday communication found in their pre-training data: underlying rationales are frequently left implicit (unstated). To address this challenge, we introduce RATIONALYST, a model for process-supervision of reasoning based on pre-training on a vast collection of rationale annotations extracted from unlabeled data. We extract 79k rationales from web-scale unlabelled dataset (the Pile) and a combination of reasoning datasets with minimal human intervention. This web-scale pre-training for reasoning allows RATIONALYST to consistently generalize across diverse reasoning tasks, including mathematical, commonsense, scientific, and logical reasoning. Fine-tuned from LLaMa-3-8B, RATIONALYST improves the accuracy of reasoning by an average of 3.9% on 7 representative reasoning benchmarks. It also demonstrates superior performance compared to significantly larger verifiers like GPT-4 and similarly sized models fine-tuned on matching training sets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。