arXiv:2506.12588cs.LG2025-06被引 1

指出时序链接预测评估中的三大漏洞,提醒警惕误导性结果。

Are We Really Measuring Progress? Transferring Insights from Evaluating Recommender Systems to Temporal Link Prediction

  • 发现当前评估存在采样不一致、硬负样本依赖和概率假设偏差问题。
  • 通过推荐系统类比揭示评估机制的深层缺陷,影响结果可信度。
  • 适合关注基准测试可靠性、评估方法设计的研究者参考。

近期研究质疑图学习基准的可靠性,关注任务设计、方法严谨性和数据适用性。本文聚焦时序链接预测(TLP)的评估策略,指出当前协议常受以下问题影响:(1) 采样指标不一致,(2) 依赖硬负样本以提升鲁棒性,(3) 评估指标隐含源节点基础概率相等的假设。通过实例分析与推荐系统领域的长期争议建立联系,支持上述观点。当前工作旨在系统刻画这些问题,并探索更稳健、可解释的替代评估方案。最后讨论改进TLP基准可靠性的潜在方向。

原文摘要 · Abstract (English)

Recent work has questioned the reliability of graph learning benchmarks, citing concerns around task design, methodological rigor, and data suitability. In this extended abstract, we contribute to this discussion by focusing on evaluation strategies in Temporal Link Prediction (TLP). We observe that current evaluation protocols are often affected by one or more of the following issues: (1) inconsistent sampled metrics, (2) reliance on hard negative sampling often introduced as a means to improve robustness, and (3) metrics that implicitly assume equal base probabilities across source nodes by combining predictions. We support these claims through illustrative examples and connections to longstanding concerns in the recommender systems community. Our ongoing work aims to systematically characterize these problems and explore alternatives that can lead to more robust and interpretable evaluation. We conclude with a discussion of potential directions for improving the reliability of TLP benchmarks.

时序链接评估方法基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。