厘清常识知识定义,发现主流数据集含大量非常识题
What Really is Commonsense Knowledge?
- 从三大理论框架整合出统一的常识知识定义
- 实测发现两个数据集中超半数题目不属常识知识
- 大模型在真常识题上表现明显更差,验证评测可信度问题
常识知识数据集在自然语言处理中已广泛发展,主要依赖众包人工标注。然而,关于常识推理基准的真实性存在争议:部分基准中大量题目并不涉及常识知识,这会削弱对模型真实常识推理能力的评估。该问题可能源于常识知识概念本身的模糊性。为澄清上述质疑,本研究梳理现有常识知识定义,基于三类概念界定框架进行整合,提出统一的多框架常识知识定义(称作整合定义)。随后,利用该定义对CommonsenseQA与CommonsenseQA 2.0数据集进行标注与实验,检验前述质疑。结果显示,两个数据集中均存在大量非常识知识样本,且在这些子集上大型语言模型(LLMs)的表现显著低于常识知识样本,表明当前评测体系存在严重偏差。
原文摘要 · Abstract (English)
Commonsense datasets have been well developed in Natural Language Processing, mainly through crowdsource human annotation. However, there are debates on the genuineness of commonsense reasoning benchmarks. In specific, a significant portion of instances in some commonsense benchmarks do not concern commonsense knowledge. That problem would undermine the measurement of the true commonsense reasoning ability of evaluated models. It is also suggested that the problem originated from a blurry concept of commonsense knowledge, as distinguished from other types of knowledge. To demystify all of the above claims, in this study, we survey existing definitions of commonsense knowledge, ground into the three frameworks for defining concepts, and consolidate them into a multi-framework unified definition of commonsense knowledge (so-called consolidated definition). We then use the consolidated definition for annotations and experiments on the CommonsenseQA and CommonsenseQA 2.0 datasets to examine the above claims. Our study shows that there exists a large portion of non-commonsense-knowledge instances in the two datasets, and a large performance gap on these two subsets where Large Language Models (LLMs) perform worse on commonsense-knowledge instances.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。