arXiv:2412.07273cs.LGcs.AI2024-12AAAI被引 6

提出新评估指标,揭示并缓解图神经网络的突发性错误聚集问题。

Temporal-Aware Evaluation and Learning for Temporal Graph Neural Networks

  • 设计基于波动性的评估指标,捕捉模型预测中的突发错误模式。
  • 实验证明现有模型在时间上易出现错误集中,且不同模型模式各异。
  • 将新指标用于训练,显著减少错误的突发聚集,提升稳定性。

时序图神经网络(TGNNs)旨在从动态图中建模和学习时序信息,虽已取得显著应用成效,但研究多集中于算法与系统设计,评估指标却未受足够重视。本文分析了现有评估指标的失效机制,发现其无法有效捕捉模型预测行为中的关键时序结构,特别是错误波动聚集现象——即错误在短时间内集中出现。通过数学建模与实例研究,我们揭示了该现象对算法与系统设计的重要影响。为此,提出一种新的波动性感知评估指标(波动聚类统计),用于更精细地分析模型时序性能。进一步,我们将该指标作为训练目标,以缓解时序错误的聚集。在多种TGNN模型上的实验表明:1)现有模型普遍存在错误波动聚集;2)不同机制的模型表现出不同的波动模式。所提训练目标能有效降低错误聚类,提升模型稳定性。

原文摘要 · Abstract (English)

Temporal Graph Neural Networks (TGNNs) are a family of graph neural networks designed to model and learn dynamic information from temporal graphs. Given their substantial empirical success, there is an escalating interest in TGNNs within the research community. However, the majority of these efforts have been channelled towards algorithm and system design, with the evaluation metrics receiving comparatively less attention. Effective evaluation metrics are crucial for providing detailed performance insights, particularly in the temporal domain. This paper investigates the commonly used evaluation metrics for TGNNs and illustrates the failure mechanisms of these metrics in capturing essential temporal structures in the predictive behaviour of TGNNs. We provide a mathematical formulation of existing performance metrics and utilize an instance-based study to underscore their inadequacies in identifying volatility clustering (the occurrence of emerging errors within a brief interval). This phenomenon has profound implications for both algorithm and system design in the temporal domain. To address this deficiency, we introduce a new volatility-aware evaluation metric (termed volatility cluster statistics), designed for a more refined analysis of model temporal performance. Additionally, we demonstrate how this metric can serve as a temporal-volatility-aware training objective to alleviate the clustering of temporal errors. Through comprehensive experiments on various TGNN models, we validate our analysis and the proposed approach. The empirical results offer revealing insights: 1) existing TGNNs are prone to making errors with volatility clustering, and 2) TGNNs with different mechanisms to capture temporal information exhibit distinct volatility clustering patterns. Our empirical findings demonstrate that our proposed training objective effectively reduces volatility clusters in error.

图神经网络时序建模评估指标误差分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。