arXiv:2504.08970cs.LGcs.AI2025-04被引 4

大规模评估嵌入模型在知识图谱补全中的表现,揭示现有评测的严重缺陷。

On Large-scale Evaluation of Embedding Models for Knowledge Graph Completion

  • 在真实数据集上测试四种主流模型,突破小规模数据局限
  • 发现模型性能在大小数据集间差异显著,排名与得分均不稳定
  • 指出当前评测协议高估模型能力,尤其在处理复杂关系时

知识图谱嵌入(KGE)模型广泛用于知识图谱补全,但其评估长期受限于不现实的基准。标准指标基于封闭世界假设,错误惩罚正确预测的缺失三元组,违背链接预测的根本目标。这些指标常将准确率压缩为单一数值,掩盖模型的具体优劣。主流评估协议“链接预测”假定待预测实体属性已知,这在现实中不合理。尽管存在属性预测、实体对排序、三元组分类等替代方案,却未被充分使用。此外,常用数据集或有缺陷,或规模过小,无法反映真实场景。极少研究关注中介节点的作用(对建模n-ary关系至关重要),也缺乏跨领域性能分析。本文在大规模数据集FB-CVT-REV和FB+CVT-REV上对四种代表性KGE模型进行综合评估。结果揭示关键洞见:小样本与大样本下模型性能差异显著,相对排名与绝对得分均发生改变;将n-ary关系二值化会系统性高估模型能力;现有评估协议与指标存在根本性缺陷。

原文摘要 · Abstract (English)

Knowledge graph embedding (KGE) models are extensively studied for knowledge graph completion, yet their evaluation remains constrained by unrealistic benchmarks. Standard evaluation metrics rely on the closed-world assumption, which penalizes models for correctly predicting missing triples, contradicting the fundamental goals of link prediction. These metrics often compress accuracy assessment into a single value, obscuring models' specific strengths and weaknesses. The prevailing evaluation protocol, link prediction, operates under the unrealistic assumption that an entity's properties, for which values are to be predicted, are known in advance. While alternative protocols such as property prediction, entity-pair ranking, and triple classification address some of these limitations, they remain underutilized. Moreover, commonly used datasets are either faulty or too small to reflect real-world data. Few studies examine the role of mediator nodes, which are essential for modeling n-ary relationships, or investigate model performance variation across domains. This paper conducts a comprehensive evaluation of four representative KGE models on large-scale datasets FB-CVT-REV and FB+CVT-REV. Our analysis reveals critical insights, including substantial performance variations between small and large datasets, both in relative rankings and absolute metrics, systematic overestimation of model capabilities when n-ary relations are binarized, and fundamental limitations in current evaluation protocols and metrics.

知识图谱嵌入模型评估方法大数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。