arXiv:2508.15291cs.LGcs.CL2025-08

提出新指标评估知识图谱复杂度,发现旧指标不靠谱。

Evaluating Knowledge Graph Complexity via Semantic, Spectral, and Structural Metrics for Link Prediction

  • 用语义、谱和结构指标分析知识图谱复杂度
  • 关系熵等指标与预测性能强负相关,更反映任务难度
  • 原宣称稳定的谱梯度指标在实际中表现不佳

理解数据集复杂度是评估和比较知识图谱(KG)上链接预测模型的基础。尽管累积谱梯度(CSG)被提出为一种与分类器无关的复杂度度量,声称其随类别数量增长且与下游性能相关,但尚未在知识图谱场景中验证。本文在多关系链接预测背景下,结合基于Transformer的语义嵌入,对CSG进行批判性分析。结果表明,CSG对参数设置敏感,不随类别数稳定增长,且与标准指标如均倒数排名(MRR)和Hit@1的相关性弱或不一致。为进一步分析,我们引入并基准测试了一组结构与语义复杂度度量。发现全局与局部关系模糊性(通过关系熵、节点级最大关系多样性、关系类型基数刻画)与MRR和Hit@1呈强负相关,提示其为更可靠的难度指标;而图连通性度量如平均度、度熵、PageRank和特征向量中心性则与Hit@10正相关。结果表明,CSG宣称的稳定性与泛化预测能力在链接预测中不成立,凸显需要更稳定、可解释且任务对齐的数据集复杂度度量。

原文摘要 · Abstract (English)

Understanding dataset complexity is fundamental to evaluating and comparing link prediction models on knowledge graphs (KGs). While the Cumulative Spectral Gradient (CSG) metric, derived from probabilistic divergence between classes within a spectral clustering framework, has been proposed as a classifier agnostic complexity metric purportedly scaling with class cardinality and correlating with downstream performance, it has not been evaluated in KG settings so far. In this work, we critically examine CSG in the context of multi relational link prediction, incorporating semantic representations via transformer derived embeddings. Contrary to prior claims, we find that CSG is highly sensitive to parametrisation and does not robustly scale with the number of classes. Moreover, it exhibits weak or inconsistent correlation with standard performance metrics such as Mean Reciprocal Rank (MRR) and Hit@1. To deepen the analysis, we introduce and benchmark a set of structural and semantic KG complexity metrics. Our findings reveal that global and local relational ambiguity captured via Relation Entropy, node level Maximum Relation Diversity, and Relation Type Cardinality exhibit strong inverse correlations with MRR and Hit@1, suggesting these as more faithful indicators of task difficulty. Conversely, graph connectivity measures such as Average Degree, Degree Entropy, PageRank, and Eigenvector Centrality correlate positively with Hit@10. Our results demonstrate that CSGs purported stability and generalization predictive power fail to hold in link prediction settings and underscore the need for more stable, interpretable, and task-aligned measures of dataset complexity in knowledge driven learning.

知识图谱复杂度评估链接预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。