arXiv:2512.06296cs.AI2025-12被引 2

提出新评估框架PROBE,更准确衡量知识图谱补全模型的预测精度与公平性。

How Sharp and Bias-Robust is a Model? Dual Evaluation Perspectives on Knowledge Graph Completion

  • 引入预测锐度与流行度偏差鲁棒性双视角评估
  • 在真实数据集上发现传统指标存在严重偏差
  • 适合关注模型可靠性与公平性的研究者

知识图谱补全(KGC)旨在从已知事实中预测缺失条目。尽管已有众多KGC模型,其评估方法仍不充分。本文指出现有指标忽略两个关键维度:(A1) 预测锐度——对单个预测严格程度的衡量;(A2) 流行度偏差鲁棒性——对低流行度实体的预测能力。为此,我们提出新评估框架PROBE,包含基于所需预测锐度估计得分的排名变换器(RT)和以流行度感知方式聚合得分的排名聚合器(RA)。在真实世界KG上的实验表明,现有指标常高估或低估模型性能,而PROBE能提供更全面、可靠的评估结果。

原文摘要 · Abstract (English)

Knowledge graph completion (KGC) aims to predict missing facts from the observed KG. While a number of KGC models have been studied, the evaluation of KGC still remain underexplored. In this paper, we observe that existing metrics overlook two key perspectives for KGC evaluation: (A1) predictive sharpness -- the degree of strictness in evaluating an individual prediction, and (A2) popularity-bias robustness -- the ability to predict low-popularity entities. Toward reflecting both perspectives, we propose a novel evaluation framework (PROBE), which consists of a rank transformer (RT) estimating the score of each prediction based on a required level of predictive sharpness and a rank aggregator (RA) aggregating all the scores in a popularity-aware manner. Experiments on real-world KGs reveal that existing metrics tend to over- or under-estimate the accuracy of KGC models, whereas PROBE yields a comprehensive understanding of KGC models and reliable evaluation results.

知识图谱评估方法模型公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。