arXiv:2409.07638cs.AIcs.CL2024-09被引 18

GPT-4在简单任务中表现受提示和输入细节影响极大,揭示评估误区。

Can We Count on LLMs? The Fixed-Effect Fallacy and Claims of GPT-4 Capabilities

  • 设计确定性任务测试GPT-4,控制变量观察性能变化
  • 相同任务下,提示语和输入元素频率导致准确率差异超统计波动范围
  • 提醒研究者警惕语言表达微调对模型表现的误导性影响

本文探讨大语言模型能力评估问题。通过在多个确定性任务中测量GPT-4的表现,每个任务包含基础计算并以大型定义明确的输入集合为参数(如列表计数、两个k位数相乘等)。每项任务在多种条件下进行足够多的试验,以检测统计上显著的差异。研究发现,任务提示的细微修改或输入集合的变化,会导致性能差异远超采样误差范围。例如,在列表计数任务中,准确率不仅随列表长度变化,还受列表组成(即被计数对象)和元素频率影响——当某元素占列表约50%时与占约70%时表现明显不同。结论表明,量化大模型能力易陷入语言作为固定效应的谬误,即错误地将有限观测泛化到更广泛情境。由此产生的直觉往往基于人类交互经验,但对输入微小调整是否应无影响的判断极不可靠。

原文摘要 · Abstract (English)

In this paper we explore evaluation of LLM capabilities. We present measurements of GPT-4 performance on several deterministic tasks; each task involves a basic calculation and takes as input parameter some element drawn from a large well-defined population (e.g., count elements in a list, multiply two k-digit numbers, etc). We examine several conditions per-task and perform enough trials so that statistically significant differences can be detected. This allows us to investigate the sensitivity of task-accuracy both to query phrasing and input parameter population. We find that seemingly trivial modifications in the task-prompt or input population can yield differences far larger than can be explained by sampling effects. For example, performance on a simple list-counting task varies with query-phrasing and list-length, but also with list composition (i.e., the thing-to-be-counted) and object frequency (e.g., success when an element accounts for $\approx$ 50\% of a list is different from when it accounts for $\approx$ 70\% etc). We conclude that efforts to quantify LLM capabilities easily succumb to the language-as-fixed-effect fallacy, where experimental observations are improperly generalized beyond what the data supports. A consequence appears to be that intuitions that have been formed based on interactions with humans form a very unreliable guide as to which input modifications should ``make no difference'' to LLM performance.

LLM评估提示工程偏差分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。