提出新方法提升图文图数据的高效学习与可读性。
Semi-Supervised Text-Attributed Graph Distillation

- 用双路径编码融合图文本信息,自训练生成可靠伪标签。
- 基于最优传输距离压缩图结构,保持下游任务性能。
- 生成可读文本摘要,适合大模型应用的图文图分析。
图文图(TAGs)作为融合图拓扑与丰富文本语义的表达模型日益重要。现有基于图神经网络的表示学习方法在处理大型语言模型时面临严重可扩展性瓶颈。数据蒸馏虽为数据驱动的解决方案,但现有方法未能捕捉图与文本模态间的复杂交互,难以应对半监督场景中的标签稀缺问题,且无法生成下游大模型任务所需的可读文本属性。为此,我们提出 extsc{algo},一种基于最优传输距离(WSD)的统一半监督框架。基于真实TAGs的实证发现, extsc{algo} 引入图-文协同编码模块,在协同自训练机制下使用双路径编码器(图感知与图无关)挖掘可靠伪标签,并融合互补的图-文特征。此外,我们设计了理论支撑的WSD图剪枝算法与低成本的大模型文本生成模块,通过聚类关键词提取生成连贯、可读的节点摘要。在基准数据集上的大量实验表明, extsc{algo} 在图神经网络和大模型下游任务中均实现了最优的性能-压缩权衡,支持高效、有效的图文图学习与分析。
原文摘要 · Abstract (English)
{\em Text-Attributed Graphs} (TAGs) have emerged as an expressive data model for integrating graph topology with rich textual semantics. Existing representation learning methods over TAGs suffer from severe scalability bottlenecks, particularly together with {\em Large Language Models} (LLMs). While data distillation offers a promising data-centric solution, existing methods fail to capture the complex interplay between graph and text modalities, struggle with the label scarcity inherent in semi-supervised settings, and lack the ability to produce the human-readable textual attributes required for downstream LLM-based tasks. To address these challenges, we propose \algo{}, a unified semi-supervised framework guided by the {\em Wasserstein Distance} (WSD). Grounded in our empirical findings on real TAGs, \algo{} introduces a graph-text collaborative encoding module that utilizes dual-pathway encoders (graph-aware and -free) within a collaborative self-training scheme to harvest reliable pseudo-labels and fuse complementary graph-text features. Furthermore, we develop a theoretically grounded WSD-based graph sketching algorithm and a cost-effective LLM text synthesis module, which leverages cluster-based keyword extraction to generate coherent, human-readable summaries for condensed nodes. Extensive experiments on benchmark datasets demonstrate that \algo{} achieves a state-of-the-art performance-compression trade-off in terms of both GNN- and LLM-based downstream tasks, enabling effective and efficient TAG learning or analytics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。