arXiv:2502.11393cs.CL2025-02ACL被引 7

构建首个双语常识推理鲁棒性评测基准,揭示大模型在常识理解上的脆弱性。

HellaSwag-Pro: A Large-Scale Bilingual Benchmark for Evaluating the Robustness of LLMs in Commonsense Reasoning

  • 设计七类问题变体,构建1.12万例双语评测数据集
  • 41个大模型测试显示常识推理普遍不鲁棒,中英文表现差异明显
  • 提供细粒度标注的中文数据集,助力跨语言认知研究

大语言模型在常识推理方面表现出显著能力,但问题表述的细微变化常导致错误回答。这些模型是真正理解常识,还是仅记忆表达模式?为探究此问题,我们首次开展大规模的常识推理鲁棒性评估。提出HellaSwag-Pro,一个包含11,200个案例的双语基准,通过设计并整合七类问题变体构建。为构建该基准,我们提出两阶段方法,开发了包含12,000个实例、覆盖56个类别的精细标注中文数据集HellaSwag。在41个代表性大模型上进行广泛实验,发现这些模型在常识推理中远未达到鲁棒性;且其表现随测试语言不同而显著变化。本工作建立高质量评估基准,广泛实验为社区提供了关于大模型常识推理的重要洞见。

原文摘要 · Abstract (English)

Large language models (LLMs) have shown remarkable capabilities in commonsense reasoning; however, some variations in questions can trigger incorrect responses. Do these models truly understand commonsense knowledge, or just memorize expression patterns? To investigate this question, we present the first extensive robustness evaluation of LLMs in commonsense reasoning. We introduce HellaSwag-Pro, a large-scale bilingual benchmark consisting of 11,200 cases, by designing and compiling seven types of question variants. To construct this benchmark, we propose a two-stage method to develop Chinese HellaSwag, a finely annotated dataset comprising 12,000 instances across 56 categories. We conduct extensive experiments on 41 representative LLMs, revealing that these LLMs are far from robust in commonsense reasoning. Furthermore, this robustness varies depending on the language in which the LLM is tested. This work establishes a high-quality evaluation benchmark, with extensive experiments offering valuable insights to the community in commonsense reasoning for LLMs.

常识推理大模型评测双语基准鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。