arXiv:2605.13153cs.AI2026-05中稿 · IJCAI

提出关注稀有关键事件的评估新框架,让模型真正展现推理能力

Strikingness-Aware Evaluation for Temporal Knowledge Graph Reasoning

论文配图:Strikingness-Aware Evaluation for Temporal Knowledge Graph Reasoning
图 1 · 摘自论文原文
  • 基于时序规则计算事件罕见度,量化其重要性
  • 高罕见度事件上模型表现显著下降,揭示真实推理挑战
  • 适合追求精准评估与真实推理能力的研究者

时序知识图谱推理(TKGR)旨在从历史数据中推断缺失(尤其是未来)事件。现有评估方法对所有事件一视同仁,忽略了多数为平凡重复,导致对模型推理能力的过高估计。因此,应突出预测稀有且关键事件的能力。为此,我们提出一种感知罕见度的评估框架,引入基于规则的罕见度度量框架(RSMF),通过比较事件预期发生频率与同类型事件的基准值,量化其罕见程度,并将该值作为权重融入加权MRR和Hits@k等指标。在四个TKG基准上的实验表明:1)所有代表性模型在事件罕见度越高时性能越差;2)路径方法在低罕见度事件上表现更优,表征方法在高罕见度事件上更胜一筹;3)设计的集成方法的优势源于对平凡事件的拟合,而非推理提升。该框架提供更严谨的评估方式,推动研究聚焦于预测关键事件。

原文摘要 · Abstract (English)

Temporal Knowledge Graph Reasoning (TKGR) aims at inferring missing (especially future) events from historical data. Current evaluation in TKGR uniformly weights all events, ignoring that most are trivial repetitions, which overestimate the true reasoning ability. Therefore, the rare outstanding events, whose prediction demands deeper reasoning, should be distinguished and emphasized. To this end, we propose a strikingness-aware evaluation framework, which introduces a rule-based strikingness measuring framework (RSMF) to quantify event strikingness by comparing its expected occurrence with peer events derived from temporal rules. Strikingness is then integrated as a weighting factor into metrics like weighted MRR and Hits@k. Experiments on four TKG benchmarks reveal: 1) All representative models perform worse as event strikingness increases, 2) Path-based methods excel on low-strikingness events and representation-based ones on high-strikingness events, 3) We design an ensemble method whose gains stem from fitting trivial events rather than reasoning improvement. Our framework provides a more rigorous evaluation, refocusing the field on predicting outstanding events.

知识图谱时序推理评估方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。