无需微调,用大模型自验证推理链,提升零样本推理能力
Zero-Shot Verification-guided Chain of Thoughts
- 设计零样本提示COT STEP,分解推理步骤
- 构建两个零样本验证器,准确识别推理链对错
- 适用于数学与常识推理,无需额外标注数据
以往研究证明了思维链(COT)提示和验证器在引导大语言模型推理方面的有效性。但多数方法依赖微调的验证器或人工设计的少样本示例。本文聚焦于在完全零样本场景下,通过COT提示实现大模型对自身生成推理步骤的自验证。为此,我们设计了新的零样本提示COT STEP,用于辅助零样本推理步骤分解,并提出了两种新的零样本验证器提示。我们评估了验证器对推理链正确性的分类能力,并探索了不同方式利用验证分数来指导多种数学与常识推理任务中的推理过程,使用了不同大模型进行测试。
原文摘要 · Abstract (English)
Previous works have demonstrated the effectiveness of Chain-of-Thought (COT) prompts and verifiers in guiding Large Language Models (LLMs) through the space of reasoning. However, most such studies either use a fine-tuned verifier or rely on manually handcrafted few-shot examples. In contrast, in this paper, we focus on LLM-based self-verification of self-generated reasoning steps via COT prompts in a completely zero-shot regime. To explore this setting, we design a new zero-shot prompt, which we call COT STEP, to aid zero-shot decomposition of reasoning steps and design two new zero-shot prompts for LLM-based verifiers. We evaluate the verifiers' ability to classify the correctness of reasoning chains and explore different ways to use verifier scores in guiding reasoning for various mathematical and commonsense reasoning tasks with different LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。