arXiv:2606.08921cs.LG2026-06

提出新评估框架,让知识图谱补全模型表现更真实可信。

Generalized Rank-based Evaluation for Knowledge Graph Completion: Perspectives, Framework, and Analyses

论文配图:Generalized Rank-based Evaluation for Knowledge Graph Completion: Perspectives, Framework, and Analyses
图 1 · 摘自论文原文
  • 设计可调锐度与抗流行度偏见的评估框架
  • 在六大数据集上验证比现有方法更稳定可靠
  • 适合需要公平比较模型性能的研究者使用

知识图谱补全(KGC)旨在从观测到的知识图谱中预测缺失事实,在药物发现、推荐系统和检索增强生成等实际应用中至关重要。尽管已有大量KGC模型被提出,但其评估方法仍不充分,严重影响模型性能的可靠判断和实际应用中的选型。本文指出两个被忽视的评估视角:(P1)预测锐度,(P2)流行度偏见鲁棒性。为此提出通用评估框架PROBE,包含评分变换器(RT)控制预测锐度,以及排名聚合器(RA)控制流行度偏见鲁棒性。理论分析表明,PROBE满足六个可靠评估的关键性质,而现有指标均不满足。特别地,在开放世界知识图谱下,即使仅观察部分事实,评估指标也应保持模型相对性能的一致性。实验显示,现有指标可能因评估视角不同而高估或低估模型表现,而PROBE能提供更全面、灵活且一致的评估结果,适用于六种真实世界知识图谱与六种主流KGC模型。

原文摘要 · Abstract (English)

Knowledge graph completion (KGC) aims to predict missing facts from an observed knowledge graph (KG), playing a crucial role in a wide range of real-world applications such as drug discovery, recommender systems, and retrieval-augmented generation (RAG). Although numerous KGC models have been proposed, the evaluation of KGC remains underexplored, despite its critical role in reliably assessing model performance and selecting appropriate models for real-world applications. In this paper, we introduce two important perspectives for KGC evaluation that are overlooked by existing evaluation metrics, (P1) predictive sharpness and (P2) popularity-bias robustness. To address both perspectives, we propose a generalized evaluation framework, PROBE, which consists of a rank transformer (RT) that estimates the score of each prediction based on a desired level of predictive sharpness and a rank aggregator (RA) that determines the final evaluation score by aggregating all prediction scores according to a desired level of popularity-bias robustness. We theoretically analyze PROBE by defining six key properties for reliable KGC evaluation and prove that PROBE satisfies all the properties, while existing metrics fail to satisfy some. In particular, due to the open-world nature of KGs, an evaluation metric should preserve the relative performance of KGC models even when only incomplete facts are observed. We show that PROBE better maintains such consistency, providing a more reliable estimate of intrinsic model performance than existing metrics. Extensive experiments with six KGC models on six real-world KGs reveal that existing metrics may over- or under-estimate model performance depending on different evaluation perspectives, whereas PROBE enables a more comprehensive, flexible, and consistent evaluation of KGC models.

知识图谱评估方法模型对比

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。