无需训练即可高效压缩异构图,提升模型训练速度与质量
Training-free Heterogeneous Graph Condensation via Data Selection
- 将异构图压缩转为数据选择问题,基于元路径筛选关键节点
- 提出无训练的压缩策略,显著降低计算耗时并保持高质量图结构
- 适合需要快速部署异构图模型的研究者和工业应用
大规模异构图的高效训练在实际应用中至关重要。现有方法多通过简化模型来降低资源与时间开销,却忽视了从数据层面简化大规模异构图的重要性。为此,HGCond首次引入异构图压缩(GC),生成小规模凝聚图以支持高效训练。然而,其存在两大缺陷:一是效果不足,过度依赖简单中继模型,限制了异构图神经网络(HGNN)在灵活压缩比下的表现与泛化能力;二是效率低下,沿用针对同构图设计的复杂优化范式,导致压缩过程耗时。针对上述挑战,本文提出首个无需训练的异构图压缩方法FreeHGC,实现高效且高质量的异构图凝聚。具体而言,将异构图压缩重构为数据选择问题,提供评估与压缩代表性节点与边的新视角。通过利用丰富的元路径(meta-paths),设计新的高质量异构数据选择准则以筛选目标类型节点。同时,提出两种无需训练的异构图压缩策略,有效凝聚与合成其他类型节点。
原文摘要 · Abstract (English)
Efficient training of large-scale heterogeneous graphs is of paramount importance in real-world applications. However, existing approaches typically explore simplified models to mitigate resource and time overhead, neglecting the crucial aspect of simplifying large-scale heterogeneous graphs from the data-centric perspective. Addressing this gap, HGCond introduces graph condensation (GC) in heterogeneous graphs and generates a small condensed graph for efficient model training. Despite its efficacy in graph generation, HGCond encounters two significant limitations. The first is low effectiveness, HGCond excessively relies on the simplest relay model for the condensation procedure, which restricts the ability to exert powerful Heterogeneous Graph Neural Networks (HGNNs) with flexible condensation ratio and limits the generalization ability. The second is low efficiency, HGCond follows the existing GC methods designed for homogeneous graphs and leverages the sophisticated optimization paradigm, resulting in a time-consuming condensing procedure. In light of these challenges, we present the first Training \underline{Free} Heterogeneous Graph Condensation method, termed FreeHGC, facilitating both efficient and high-quality generation of heterogeneous condensed graphs. Specifically, we reformulate the heterogeneous graph condensation problem as a data selection issue, offering a new perspective for assessing and condensing representative nodes and edges in the heterogeneous graphs. By leveraging rich meta-paths, we introduce a new, high-quality heterogeneous data selection criterion to select target-type nodes. Furthermore, two training-free condensation strategies for heterogeneous graphs are designed to condense and synthesize other-types nodes effectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。