arXiv:2412.08937cs.LGcs.CL2024-12被引 7

构建跨领域多尺度文本属性异构图数据集,推动真实场景下图学习模型评估

Multi-Scale Heterogeneous Text-Attributed Graph Datasets From Diverse Domains

  • 构建覆盖电影、学术、专利等多领域的异构文本图数据集
  • 数据集涵盖多年跨度,支持不同规模和复杂度的模型测试
  • 开源全部数据与代码,助力可复现的图学习研究

异构文本属性图(HTAG)在多种领域中广泛应用,其特征是不同类型的实体不仅带有文本信息,还通过多样关系连接。然而,当前文本属性图学习研究主要聚焦于同质图(单一节点和边类型),缺乏对HTAG的有效评估数据集。主要原因在于缺少包含原始文本内容且覆盖多个领域、不同规模的综合性数据集。为此,我们构建了一组具有挑战性且多样化的基准数据集,用于真实场景下对HTAG上机器学习模型的可复现评估。这些数据集具有多尺度特性,时间跨度达数年,涵盖电影、社区问答、学术、文献和专利网络等多个领域。我们在这些数据集上对多种图神经网络进行了基准实验。所有源数据、数据集构建代码、处理后的HTAG、数据加载器、基准代码及评估设置均已公开发布于GitHub和Hugging Face。

原文摘要 · Abstract (English)

Heterogeneous Text-Attributed Graphs (HTAGs), where different types of entities are not only associated with texts but also connected by diverse relationships, have gained widespread popularity and application across various domains. However, current research on text-attributed graph learning predominantly focuses on homogeneous graphs, which feature a single node and edge type, thus leaving a gap in understanding how methods perform on HTAGs. One crucial reason is the lack of comprehensive HTAG datasets that offer original textual content and span multiple domains of varying sizes. To this end, we introduce a collection of challenging and diverse benchmark datasets for realistic and reproducible evaluation of machine learning models on HTAGs. Our HTAG datasets are multi-scale, span years in duration, and cover a wide range of domains, including movie, community question answering, academic, literature, and patent networks. We further conduct benchmark experiments on these datasets with various graph neural networks. All source data, dataset construction codes, processed HTAGs, data loaders, benchmark codes, and evaluation setup are publicly available at GitHub and Hugging Face.

异构图文本属性数据集图神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。