聚焦有向图数据质量,推动机器学习从模型为中心转向数据为中心。
Towards Data-centric Machine Learning on Directed Graphs: a Survey
- 构建有向图学习的新分类体系,强调数据表示的重要性。
- 揭示有向图结构对模型性能的关键影响,优于传统无向简化。
- 覆盖10多个领域应用,适合图神经网络与数据建模研究者参考。
近年来,图神经网络(GNN)在处理结构化数据方面取得显著进展。然而,多数方法仍采用以模型为中心的范式,将图简化为无向形式并侧重模型设计,这在真实场景中因信息丢失和模型优化困难而受限。因此,研究趋势正转向数据为中心的方法,注重提升图的质量与表示能力。自然结构数据可衍生出异质图、超图和有向图等类型,其中有向图能有效建模因果关系,在拓扑系统中具有独特优势。近年来,有向GNN受到广泛关注,但系统性综述仍不足。本文旨在提供有向图学习的全面回顾,尤其从数据为中心视角出发,提出新的研究分类体系,并重新审视现有方法,强调理解与改进数据表示的重要性。研究表明,深入理解有向图及其质量对模型性能至关重要。此外,本文还探讨了有向GNN在10余个领域的多样化应用,展示其广泛适用性。最后,识别该领域的关键机遇与挑战,为未来研究提供指引。
原文摘要 · Abstract (English)
In recent years, Graph Neural Networks (GNNs) have made significant advances in processing structured data. However, most of them primarily adopted a model-centric approach, which simplifies graphs by converting them into undirected formats and emphasizes model designs. This approach is inherently limited in real-world applications due to the unavoidable information loss in simple undirected graphs and the model optimization challenges that arise when exceeding the upper bounds of this sub-optimal data representational capacity. As a result, there has been a shift toward data-centric methods that prioritize improving graph quality and representation. Specifically, various types of graphs can be derived from naturally structured data, including heterogeneous graphs, hypergraphs, and directed graphs. Among these, directed graphs offer distinct advantages in topological systems by modeling causal relationships, and directed GNNs have been extensively studied in recent years. However, a comprehensive survey of this emerging topic is still lacking. Therefore, we aim to provide a comprehensive review of directed graph learning, with a particular focus on a data-centric perspective. Specifically, we first introduce a novel taxonomy for existing studies. Subsequently, we re-examine these methods from the data-centric perspective, with an emphasis on understanding and improving data representation. It demonstrates that a deep understanding of directed graphs and their quality plays a crucial role in model performance. Additionally, we explore the diverse applications of directed GNNs across 10+ domains, highlighting their broad applicability. Finally, we identify key opportunities and challenges within the field, offering insights that can guide future research and development in directed graph learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。