arXiv:2504.14462cs.CL2025-04被引 2

构建长尾实体常识推理数据集,揭示大模型在冷门知识上的幻觉问题。

CoLoTa: A Dataset for Entity-based Commonsense Reasoning over Long-Tail Knowledge

  • 针对冷门实体设计3300个需常识推理的问答与判断题
  • 现有大模型在长尾常识任务上错误率极高
  • 适合评估大模型和知识图谱问答系统的常识能力

大型语言模型(LLMs)虽在事实与常识知识编码及推理任务中表现优异,但在高风险场景部署时仍受幻觉和推理错误困扰。本文发现,即使OpenAI-o1等先进模型在涉及冷门、长尾实体的常识推理任务中也存在高错误率。为此,我们提出新的常识推理数据集CoLoTa,包含3,300个来自问答与声明验证任务的查询,覆盖多样化的常识推理技能。该数据集亦可作为知识图谱问答(KGQA)基准,因其答案所需知识存在于Wikidata中。但不同于仅关注事实性问题的现有KGQA基准,CoLoTa查询还需常识推理。实验表明,强效的基于LLM的KGQA方法在涉及常识推理的任务上表现严重不足。因此,CoLoTa被提议为评估(i)LLM在长尾实体上的常识推理能力与抗幻觉性,以及(ii)KGQA方法常识推理能力的新基准。

原文摘要 · Abstract (English)

The rise of Large Language Models (LLMs) has redefined the AI landscape, particularly due to their ability to encode factual and commonsense knowledge, and their outstanding performance in tasks requiring reasoning. Despite these advances, hallucinations and reasoning errors remain a significant barrier to their deployment in high-stakes settings. In this work, we observe that even the most prominent LLMs, such as OpenAI-o1, suffer from high rates of reasoning errors and hallucinations on tasks requiring commonsense reasoning over obscure, long-tail entities. To investigate this limitation, we present a new dataset for Commonsense reasoning over Long-Tail entities (CoLoTa), that consists of 3,300 queries from question answering and claim verification tasks and covers a diverse range of commonsense reasoning skills. We remark that CoLoTa can also serve as a Knowledge Graph Question Answering (KGQA) dataset since the support of knowledge required to answer its queries is present in the Wikidata knowledge graph. However, as opposed to existing KGQA benchmarks that merely focus on factoid questions, our CoLoTa queries also require commonsense reasoning. Our experiments with strong LLM-based KGQA methodologies indicate their severe inability to answer queries involving commonsense reasoning. Hence, we propose CoLoTa as a novel benchmark for assessing both (i) LLM commonsense reasoning capabilities and their robustness to hallucinations on long-tail entities and (ii) the commonsense reasoning capabilities of KGQA methods.

常识推理长尾知识大模型评测知识图谱

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。