arXiv:2503.23779cs.CLcs.AI2025-03被引 3

构建新语料WinoWhat,发现大模型常识推理能力被高估

WinoWhat: A Parallel Corpus of Paraphrased WinoGrande Sentences with Common Sense Categorization

  • 对WinoGrande数据集进行改写生成新语料WinoWhat
  • 所有模型在新语料上表现显著下降,平均准确率降低18%
  • 揭示模型在因果、物理、社会等常识类别中存在系统性短板

本研究聚焦于如何利用温格罗兰难题评估大语言模型的常识推理能力。我们评估了不同规模的生成模型在主流的WinoGrande基准上的表现。为此,我们发布了新语料WinoWhat,其中包含对WinoGrande验证集每个实例的改写版本。此外,我们在五个常识知识类别(因果、物理、社会、心理、数量)上评估模型表现,获得更细粒度的分析结果。令人意外的是,所有模型在WinoWhat上的表现均显著下降,表明当前对大模型推理能力的评估存在高估。为验证这是否由训练数据记忆导致,我们匹配基准实例与大模型训练数据,构建两套测试集。结果表明,记忆效应在WinoGrande上影响极小。

原文摘要 · Abstract (English)

In this study, we take a closer look at how Winograd schema challenges can be used to evaluate common sense reasoning in LLMs. Specifically, we evaluate generative models of different sizes on the popular WinoGrande benchmark. We release WinoWhat, a new corpus, in which each instance of the WinoGrande validation set is paraphrased. Additionally, we evaluate the performance on the challenge across five common sense knowledge categories, giving more fine-grained insights on what types of knowledge are more challenging for LLMs. Surprisingly, all models perform significantly worse on WinoWhat, implying that LLM reasoning capabilities are overestimated on WinoGrande. To verify whether this is an effect of benchmark memorization, we match benchmark instances to LLM trainingdata and create two test-suites. We observe that memorization has a minimal effect on model performance on WinoGrande.

常识推理语言模型评估基准大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。