arXiv:2412.04291cs.CL2024-12被引 3

用进化算法优化数学推理提示,准确率提升超10点

Evolutionary Pre-Prompt Optimization for Mathematical Reasoning

  • 用进化计算自动挑选最优示例构建思维链提示
  • 在GSM8K和MathQA上准确率提升超10个百分点
  • 适合想提升大模型数学推理能力的研究者

近期研究表明,大语言模型(LLMs)在仅提供少量任务相关示例时,便展现出卓越的推理能力,尤其在结合少样本学习与思维链(CoT)方法后,模型能生成更逻辑一致的结论。本文探索了为构建高效思维链预提示而进行的示例选择优化,并发现优化算法的选择——尤其是基于比较的进化计算——显著提升了效果与可行性。具体而言,得益于有限的过度探索与过拟合,进化预提示优化(EPPO)相比原始少样本方法,在GSM8K和MathQA等基准数据集上实现了超过10个绝对百分点的准确率提升。该增益在不同情境下保持一致,且在与自一致性(SC)结合时进一步放大。

原文摘要 · Abstract (English)

Recent advancements have highlighted that large language models (LLMs), when given a small set of task-specific examples, demonstrate remarkable proficiency, a capability that extends to complex reasoning tasks. In particular, the combination of few-shot learning with the chain-of-thought (CoT) approach has been pivotal in steering models towards more logically consistent conclusions [Wei et al. 2022b]. This paper explores the optimization of example selection for designing effective CoT pre-prompts and shows that the choice of the optimization algorithm, typically in favor of comparison-based methods such as evolutionary computation, significantly enhances efficacy and feasibility. Specifically, thanks to a limited exploitative and overfitted optimization, Evolutionary Pre-Prompt Optimization (EPPO) brings an improvement over the naive few-shot approach, exceeding 10 absolute points in exact match scores on benchmark datasets such as GSM8k and MathQA. These gains are consistent across various contexts and are further amplified when integrated with self-consistency (SC).

数学推理提示优化进化算法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。