arXiv:2504.07825cs.CL2025-04被引 9

HellaSwag基准存在严重有效性问题,不能可靠评估语言模型的常识推理能力。

What the HellaSwag? On the Validity of Common-Sense Reasoning Benchmarks

  • 通过分析发现该基准存在语法错误、误导性提示和选项等设计缺陷。
  • 超过65%的模型预测在替换题目后仍不变,表明评分结果不可信。
  • 研究者应慎用此基准,推荐使用修正版GoldenSwag进行评估。

常识推理是语言模型的核心能力,涵盖语言与世界理解。衡量这一能力对不同规模和应用场景的模型至关重要。目前广泛使用的基准HellaSwag存在严重构念效度问题,包括基本语法错误、大量拼写错误、误导性提示及等效正确选项。研究发现,若仅基于答案文本或用'Lorem ipsum dolor...'替代问题,超过65%的模型预测保持不变,且无法仅由数据污染解释。由于基准分数常用于研究与商业模型选型,其无效性可能导致错误决策。本文系统检验了HellaSwag的问题,并通过多种生成式语言模型验证。我们主张该基准当前状态不应再用于评估。为此,提出未来基准应满足的要求,并发布修正版GoldenSwag,以支持更可靠的常识推理评估。

原文摘要 · Abstract (English)

Common-sense reasoning is a key language model capability because it encapsulates not just specific factual knowledge but rather general language and world understanding. Measuring common-sense reasoning, therefore, is crucial for language models of different sizes and applications. One of the most widely used benchmarks for evaluating such capabilities is HellaSwag; however, in this paper, we show that it has severe construct validity issues. These issues range from basic ungrammaticality and numerous typos to misleading prompts or equally correct options. Furthermore, we show that if models are evaluated only on answer texts, or with "Lorem ipsum dolor..." instead of the question, more than 65% of model predictions remain the same, and this cannot be attributed merely to contamination. Since benchmark scores are an essential part of model selection in both research and commercial applications, these validity issues can have severe consequences. In particular, knowing that taking benchmark scores at face value is ubiquitous, inadequate evaluation leads to ill-informed decisions about models. In this paper, we thoroughly investigate critical validity issues posed by HellaSwag and illustrate them with various evaluations using generative language models of different sizes. We argue that this benchmark does not accurately measure common-sense reasoning and, therefore, should not be used for evaluation in its current state. Based on the results of our study, we propose requirements that should be met by future common-sense reasoning benchmarks. In addition, we release GoldenSwag, a corrected subset of HellaSwag, which, to our belief, facilitates acceptable common-sense reasoning evaluation.

常识推理基准评估语言模型有效性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。