动态图预测中,采样负样本会扭曲模型评估结果,应改用全实体排序。
Back to All-Entity Ranking: Sampler-Dependent Evaluation in Continuous-Time Dynamic Graphs

- 用全量目标实体替代采样负例,避免采样偏差影响评价
- 在4个数据集上,不同采样配置下模型排名顺序至少有3次变化
- 适合做动态图模型架构对比的严谨研究者
连续时间动态图中的下一个目的地预测通常将观测到的交互与采样负例进行排名。该得分同时依赖于负例分布和研究人员选择的候选数量。我们发现非均匀负例分布会改变贝叶斯最优排名,即使均匀采样的有限候选集也会导致模型排名不稳定和模块效果测量失真。随时间变化的源-目标历史成员关系以及直接使用该信息的模型操作,会将采样器的影响传递至评估分数。通过因子化评估重复与新正例对已见与未见负例的表现、基于配对历史成员关系的最小评分器及受控表示干预,我们在LastFM、MOOC、Reddit和Wikipedia上测试了六个模型。在四个数据集中的三个,至少有一对模型在统一20项指标与完整目录下的相对顺序发生变化。同一模块的效果也随候选集大小和训练目标而改变方向和幅度。这些结果表明,基于采样负例的基准测试所得出的模型优劣与消融结论均依赖于特定候选配置。全实体排名(All-entity ranking)对固定目录中的每个目标进行评估,消除负例选择自由度与采样波动,同时保留原始的CTDG评分器。因此我们建议在可枚举的固定目标目录上,以全实体排名作为动态图模型架构比较的主要依据。
原文摘要 · Abstract (English)
Next-destination prediction in continuous-time dynamic graphs (CTDGs) commonly ranks an observed interaction against sampled negative destinations. The resulting score is conditional on both the negative distribution and the number of candidates chosen by the researcher. We show that a non-uniform negative distribution changes the Bayes-optimal ranking, while even a finite candidate set drawn uniformly can destabilize model rankings and measured module effects. Time-varying source-destination history membership and model operations that use this information directly transmit the sampler's influence to the evaluation score. We examine this mechanism using a factorial evaluation of repeated and new positives against seen and unseen negatives, a minimal scorer based solely on pair-history membership, and controlled representation interventions. Across six models on LastFM, MOOC, Reddit, and Wikipedia, at least one model pair changes relative order between the expected Uniform-20 metric and the full catalog on three of the four datasets. The measured effect of the same module also changes in magnitude and direction with the candidate-set size and training objective. These results establish that model-superiority and ablation conclusions from sampled-negative benchmarks are conditional on the stated candidate configuration. All-entity ranking evaluates every destination in a fixed catalog, eliminating negative-selection freedom and sampling variation while retaining the original CTDG scorer. We therefore recommend all-entity ranking as the primary evidence for architecture comparisons on CTDG benchmarks with an enumerable, fixed destination catalog.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。