arXiv:2606.10287cs.LGcs.CL2026-06

用多准则决策方法解决知识图谱补全评估不一致问题

When Metrics Disagree: A Meta-Analysis of Knowledge-Graph-Completion Model Benchmarking

论文配图:When Metrics Disagree: A Meta-Analysis of Knowledge-Graph-Completion Model Benchmarking
图 1 · 摘自论文原文
  • 将评估视为多准则决策问题,测试七种聚合器性能
  • Z-score表现最均衡,分别推荐DualE和FMS为最优模型
  • 揭示评估稳定性与泛化能力对移除操作敏感

知识图谱补全(KGC)模型评估仍具挑战性,因常用排名指标(如MRR、Hits@k、平均秩)常在不同数据集上产生冲突的模型排序。某模型在MRR上领先,却可能在Hits@1上落后;在某一数据集表现优异,未必能泛化至其他数据集。这种碎片化阻碍了公平比较,易导致选择性报告,掩盖真实进展。本文将评估重构为多准则决策(MCDM)问题,对七种聚合器在五个测试中进行元分析:一致性、跨数据集稳定性、指标独立性、抗噪声鲁棒性及泛化能力。每项测试均通过留一模型(LOMO)和留一组(LOGO)移除方式取平均,以检验聚合器在多样模型子集中的行为可靠性。针对尾部预测(h,r,?)和关系预测(h,?,t),帕累托最优分析识别出Z-score为最均衡聚合器,分别将DualE列为尾部预测最优,FMS为关系预测最优。敏感性分析显示,一致性与稳定性基本不受移除影响,而泛化能力和独立性最为敏感。该框架有效解决评估不一致问题,为聚合器选择与模型基准测试提供实证指导。

原文摘要 · Abstract (English)

Evaluating Knowledge Graph Completion (KGC) models remains challenging because standard assessment relies on isolated rank-based metrics such as MRR, Hits$@$k, and Mean Rank, which often produce conflicting model orderings across datasets. A model that leads on MRR may trail on Hits@1, and strong performance on one dataset may not generalize to another. This fragmentation hinders comparison, enables selective reporting, and obscures real progress. We reframe KGC evaluation as a Multi-Criteria Decision-Making (MCDM) problem and present a meta-analysis of seven aggregators across five tests: consistency, cross-dataset stability, metric independence, robustness under noise, and generalizability. Each test is averaged over leave-one-model-out (LOMO) and leave-one-group-out (LOGO) removals so that reliability reflects aggregator behavior across diverse model subsets. Across tail $(h,r,?)$ and relation $(h,?,t)$ prediction, Pareto-optimal analysis identifies Z-score as the most balanced aggregator, which ranks DualE highest for tail prediction and FMS (Flow-Modulated Scoring) highest for relation prediction. A test-sensitivity analysis using the same removals shows that consistency and stability are largely removal-invariant, while generalizability and independence are the most sensitive. The framework resolves evaluation inconsistencies and offers evidence-based guidance for aggregator selection and model benchmarking in KGC.

知识图谱评估方法多准则决策模型对比

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。