arXiv:2409.20303cs.CLcs.AI2024-09被引 12

测试多种提示技巧后发现,大模型行为研究存在可复现性危机。

A Looming Replication Crisis in Evaluating Behavior in Language Models? Evidence and Solutions

  • 用多个基准测试验证五种提示工程方法的效果
  • 多数方法在各模型上均无显著差异,结果不可靠
  • 建议建立严谨评估框架以提升研究可信度

随着大语言模型(LLMs)广泛应用于各类日常场景,其行为研究迅速增长。然而,由于该领域尚处初期,缺乏明确的方法论规范,导致研究结果的可复现性和泛化能力存疑。本文通过一系列复制实验,检验了链式思维、情绪提示、专家提示、降级提示和重读等五类旨在影响模型推理能力的提示工程技巧。实验覆盖GPT-3.5、GPT-4o、Gemini 1.5 Pro、Claude 3 Opus、Llama 3-8B和Llama 3-70B六款模型,采用CommonsenseQA、CRT、NumGLUE、ScienceQA和StrategyQA等手动双检子集进行评估。结果显示,几乎所有技术在统计上均未表现出显著差异,暴露出此前研究中存在多重方法学缺陷。为此,我们提出前瞻性的解决方案:发展稳健的评估方法、构建可靠基准、设计严谨实验框架,以确保对模型输出的准确与可信评估。

原文摘要 · Abstract (English)

In an era where large language models (LLMs) are increasingly integrated into a wide range of everyday applications, research into these models' behavior has surged. However, due to the novelty of the field, clear methodological guidelines are lacking. This raises concerns about the replicability and generalizability of insights gained from research on LLM behavior. In this study, we discuss the potential risk of a replication crisis and support our concerns with a series of replication experiments focused on prompt engineering techniques purported to influence reasoning abilities in LLMs. We tested GPT-3.5, GPT-4o, Gemini 1.5 Pro, Claude 3 Opus, Llama 3-8B, and Llama 3-70B, on the chain-of-thought, EmotionPrompting, ExpertPrompting, Sandbagging, as well as Re-Reading prompt engineering techniques, using manually double-checked subsets of reasoning benchmarks including CommonsenseQA, CRT, NumGLUE, ScienceQA, and StrategyQA. Our findings reveal a general lack of statistically significant differences across nearly all techniques tested, highlighting, among others, several methodological weaknesses in previous research. We propose a forward-looking approach that includes developing robust methodologies for evaluating LLMs, establishing sound benchmarks, and designing rigorous experimental frameworks to ensure accurate and reliable assessments of model outputs.

大模型评估提示工程可复现性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。