arXiv:2412.12157cs.CLcs.AI2024-12被引 9

提出新方法提升大模型数学推理中的上下文学习效果

What Makes In-context Learning Effective for Mathematical Reasoning: A Theoretical Analysis

  • 从语义相似性和推理稳定性出发,理论分析上下文示范影响
  • 所提方法在三个基准上均实现稳定提升,超越现有方案
  • 适合希望提升大模型少样本数学推理能力的研究者

由于上下文学习能力,大语言模型(LLMs)在多个数学推理基准上表现出色。然而,我们发现少量示范有时会导致性能下降,其对推理能力的提升效果仍不可靠。为此,本文从理论上分析上下文示范对LLMs推理性能的影响。证明了推理有效性(以经验预测损失衡量)可被一个面向LLM的语义相似性与示范的推理稳定性所界定,该结论适用于单次和少次情形。基于此,提出一种简单、通用且低复杂度的示范选择方法LMS3,能自适应为不同LLM筛选最相关样本,并引入新颖的示范剔除机制,自动过滤不适合少次学习的样本。在三个代表性基准、两种LLM主干及多种少次设置下实验验证,LMS3表现出优越性,在所有数据集上均实现一致改进,现有方法无法达成此效果。

原文摘要 · Abstract (English)

Owing to the capability of in-context learning, large language models (LLMs) have shown impressive performance across diverse mathematical reasoning benchmarks. However, we find that few-shot demonstrations can sometimes bring negative performance and their effectiveness on LLMs' reasoning abilities remains unreliable. To this end, in this paper, we aim to theoretically analyze the impact of in-context demonstrations on LLMs' reasoning performance. We prove that the reasoning efficacy (measured by empirical prediction loss) can be bounded by a LLM-oriented semantic similarity and an inference stability of demonstrations, which is general for both one-shot and few-shot scenarios. Based on this finding, we propose a straightforward, generalizable, and low-complexity demonstration selection method named LMS3. It can adaptively facilitate to select the most pertinent samples for different LLMs and includes a novel demonstration rejection mechanism to automatically filter out samples that are unsuitable for few-shot learning. Through experiments on three representative benchmarks, two LLM backbones, and multiple few-shot settings, we verify that our LMS3 has superiority and achieves consistent improvements on all datasets, which existing methods have been unable to accomplish.

大模型数学推理上下文学习示范选择

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。