构建首个高质量动态图文图生成基准,推动复杂系统建模研究
GDGB: A Benchmark for Generative Dynamic Text-Attributed Graph Learning
- 设计8个高质图文属性动态图数据集,解决以往文本质量差问题
- 提出两类新生成任务:可迁移与归纳式动态图生成,覆盖真实场景演化需求
- 开发基于大模型的多智能体框架,支持可复现的生成评估与对比
动态图文图(DyTAGs)融合结构、时间与文本属性,对建模复杂现实系统至关重要。然而现有DyTAG数据集普遍文本质量低,严重限制生成任务的语义表达能力;且以往研究多聚焦判别任务,缺乏针对生成任务的标准范式与评估协议。为此,我们提出生成式动态图文图基准GDGB,包含8个精心构建的数据集,具备高质量节点与边的文本特征,突破了先前数据集的局限。基于GDGB,我们定义两项新生成任务:传递式动态图生成(TDGG)与归纳式动态图生成(IDGG)。TDGG基于源目标节点集生成目标图,而更具挑战性的IDGG引入新节点生成,以归纳建模真实图数据的动态扩展。为实现全面评估,我们设计涵盖结构、时间与文本质量的多维指标。进一步提出GAG-General——一个面向可复现与鲁棒性评测的基于大模型的多智能体生成框架。实验表明,GDGB能有效支撑TDGG与IDGG的严谨评估,关键发现揭示结构与文本特征在生成中的协同作用。这些成果确立了GDGB作为推进生成式DyTAG研究的基础资源,并助力其在实际应用中的拓展。数据集与代码已开源于https://github.com/Lucas-PJ/GDGB-ALGO。
原文摘要 · Abstract (English)
Dynamic Text-Attributed Graphs (DyTAGs), which intricately integrate structural, temporal, and textual attributes, are crucial for modeling complex real-world systems. However, most existing DyTAG datasets exhibit poor textual quality, which severely limits their utility for generative DyTAG tasks requiring semantically rich inputs. Additionally, prior work mainly focuses on discriminative tasks on DyTAGs, resulting in a lack of standardized task formulations and evaluation protocols tailored for DyTAG generation. To address these critical issues, we propose Generative DyTAG Benchmark (GDGB), which comprises eight meticulously curated DyTAG datasets with high-quality textual features for both nodes and edges, overcoming limitations of prior datasets. Building on GDGB, we define two novel DyTAG generation tasks: Transductive Dynamic Graph Generation (TDGG) and Inductive Dynamic Graph Generation (IDGG). TDGG transductively generates a target DyTAG based on the given source and destination node sets, while the more challenging IDGG introduces new node generation to inductively model the dynamic expansion of real-world graph data. To enable holistic evaluation, we design multifaceted metrics that assess the structural, temporal, and textual quality of the generated DyTAGs. We further propose GAG-General, an LLM-based multi-agent generative framework tailored for reproducible and robust benchmarking of DyTAG generation. Experimental results demonstrate that GDGB enables rigorous evaluation of TDGG and IDGG, with key insights revealing the critical interplay of structural and textual features in DyTAG generation. These findings establish GDGB as a foundational resource for advancing generative DyTAG research and unlocking further practical applications in DyTAG generation. The dataset and source code are available at https://github.com/Lucas-PJ/GDGB-ALGO.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。