arXiv:2609.08226cs.AI2026-09

首个同时评估结构与语义演化的时序图基准,揭示模型在两类任务上的能力差异。

TTGBench: Benchmarking Topological Evolution and Semantic Drift in Text-attributed Temporal Graphs

论文配图:TTGBench: Benchmarking Topological Evolution and Semantic Drift in Text-attributed Temporal Graphs
图 1 · 摘自论文原文
  • 构建六大数据集,支持多分类多标签语义追踪,解决传统数据重复问题。
  • 17种主流方法对比显示:图神经网络擅长结构预测,大模型更擅长语义跟踪。
  • 首次系统评估时序图中语义漂移,适合研究动态图与多模态学习的学者。

时序图学习建模动态系统的演化过程,其中结构关系与语义状态随时间变化。然而,现有基准主要关注通过时序链接预测(TLP)的结构演化,对语义演化的支持有限。尽管有时包含时序节点分类(TNC),但通常局限于简单的二分类设定,无法捕捉真实的语义漂移。此外,常用数据集存在高链接重复问题,导致性能估计虚高,掩盖了模型真实能力。为解决上述局限,我们提出 extbf{TTGBench},一个联合评估结构与语义演化的新型基准。该基准包含六个真实世界、文本丰富的数据集,具有 extit{双重波动性}(Dual Volatility),可实现严格公平的模型评估。尤为关键的是,它是首个支持多分类和多标签 TNC 的基准,填补了评估时序语义漂移的关键空白。我们在 17 种前沿方法(涵盖时序图神经网络 TGNN 与大语言模型 LLM 基于范式)上进行全面评估。结果揭示两大范式间存在明显的能力鸿沟:基于 TGNN 的方法在结构预测上表现优异,但在语义追踪上失败;而基于 LLM 的预测器则呈现相反趋势。通过深入分析,我们揭示其根本局限,并为开发更全面的时序图模型提供洞见。

原文摘要 · Abstract (English)

Temporal graph learning models the evolution of dynamic systems, where both structural interactions and semantic states change over time. However, existing benchmarks primarily emphasize structural evolution via temporal link prediction (TLP), while support for semantic evolution remains limited. Although temporal node classification (TNC) is sometimes included, it is typically restricted to simplistic binary settings that fail to capture realistic semantic drift. Moreover, commonly used datasets exhibit high link repetition, leading to inflated performance estimates and obscuring true model capability. To address these limitations, we introduce \textbf{TTGBench}, a new benchmark that jointly evaluates structural and semantic evolution. TTGBench comprises six real-world, text-rich datasets characterized by \emph{Dual Volatility}, enabling rigorous and fair evaluation of existing models. Notably, it is the first benchmark to support both multi-class and multi-label TNC, filling a critical gap in evaluating temporal semantic drift. We conduct a comprehensive evaluation of 17 state-of-the-art methods across Temporal Graph Neural Networks (TGNNs) and Large Language Model (LLM)-based paradigms. The results reveal a clear \emph{capability divide} between the two paradigms: TGNN-based methods excel at structural prediction but fail at semantic tracking, whereas LLM-based predictors show the opposite trend. Through in-depth analysis, we uncover their fundamental limitations and provide insights for developing more comprehensive temporal graph models.

时序图语义漂移基准测试多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。