arXiv:2505.23461cs.CL2025-05ACL被引 4

构建含事实知识的无答案问题数据集,评估大模型真实知识利用能力

UAQFact: Evaluating Factual Knowledge Utilization of LLMs on Unanswerable Questions

  • 基于知识图谱构建双语无答案问题数据集,附带辅助事实知识
  • 多模型测试显示即使有知识也难准确判断无答案问题
  • 外部知识可提升表现,但模型仍无法充分使用导致误答

处理无答案问题(UAQ)对大模型至关重要,能避免复杂场景下的误导性回复。现有评估数据集缺乏事实知识支持,难以衡量模型在处理无答案问题时对内部或外部知识的利用能力。为此,我们提出新数据集UAQFact,一个基于知识图谱构建的双语数据集,包含辅助事实知识。基于此,我们定义两项新任务:评估模型利用内部知识与外部知识的能力。实验结果表明,多个大模型系列在UAQFact上表现不佳,即使拥有相关知识也无法稳定识别无答案问题。此外,引入外部知识虽有一定提升,但模型仍未能充分使用,仍可能产生错误回答。

原文摘要 · Abstract (English)

Handling unanswerable questions (UAQ) is crucial for LLMs, as it helps prevent misleading responses in complex situations. While previous studies have built several datasets to assess LLMs' performance on UAQ, these datasets lack factual knowledge support, which limits the evaluation of LLMs' ability to utilize their factual knowledge when handling UAQ. To address the limitation, we introduce a new unanswerable question dataset UAQFact, a bilingual dataset with auxiliary factual knowledge created from a Knowledge Graph. Based on UAQFact, we further define two new tasks to measure LLMs' ability to utilize internal and external factual knowledge, respectively. Our experimental results across multiple LLM series show that UAQFact presents significant challenges, as LLMs do not consistently perform well even when they have factual knowledge stored. Additionally, we find that incorporating external knowledge may enhance performance, but LLMs still cannot make full use of the knowledge which may result in incorrect responses.

大模型评估知识利用无答案问题

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。