提出检测大模型与知识图谱间语义分歧的基准,揭示语言理解差异。
A Benchmark for the Detection of Metalinguistic Disagreements between LLMs and Knowledge Graphs
- 构建基于T-REx数据集的评测基准,区分事实与语义分歧
- 发现大模型与知识图谱存在显著语义层面的表达差异
- 适合知识图谱构建与大模型评估方向的研究者参考
在支持知识图谱构建的任务中,常以知识图谱为标准答案评估大语言模型(LLMs)的事实提取准确率。此类评估默认错误源于事实分歧,但人类对话中常出现语义分歧——即参与者对语言表达意义的理解不同,而非事实本身。鉴于自然语言处理与生成的复杂性,我们探究大模型与知识图谱之间是否存在语义分歧。基于T-REx知识对齐数据集的分析,我们假设此类分歧确实存在,并具有知识图谱工程的实际意义。为此,我们提出一个用于评估大模型与知识图谱间事实与语义分歧检测能力的基准,其初步实现已开源于Github。
原文摘要 · Abstract (English)
Evaluating large language models (LLMs) for tasks like fact extraction in support of knowledge graph construction frequently involves computing accuracy metrics using a ground truth benchmark based on a knowledge graph (KG). These evaluations assume that errors represent factual disagreements. However, human discourse frequently features metalinguistic disagreement, where agents differ not on facts but on the meaning of the language used to express them. Given the complexity of natural language processing and generation using LLMs, we ask: do metalinguistic disagreements occur between LLMs and KGs? Based on an investigation using the T-REx knowledge alignment dataset, we hypothesize that metalinguistic disagreement does in fact occur between LLMs and KGs, with potential relevance for the practice of knowledge graph engineering. We propose a benchmark for evaluating the detection of factual and metalinguistic disagreements between LLMs and KGs. An initial proof of concept of such a benchmark is available on Github.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。