arXiv:2601.21309cs.LG2026-01AAAI

提出可迁移的图数据压缩方法,提升跨任务跨域表现

Transferable Graph Condensation from the Causal Perspective

  • 从因果不变性出发,提取图结构中的稳定特征
  • 在5个公开数据集上跨任务跨域性能提升13.41%
  • 适合需要模型迁移能力的图学习场景

图数据规模的增大显著提升了图表示学习的效果,但也带来了巨大的训练挑战。图数据压缩技术通过将大规模数据集压缩为信息丰富的小型数据集,在保持相似测试性能的同时降低计算成本。然而,现有方法严格要求下游任务与原始数据集一致,难以适用于跨任务和跨域场景。为此,本文提出一种基于因果不变性的可迁移图数据压缩方法TGCC,能够生成具有强迁移能力的压缩数据集。首先,利用因果干预从图的空间域中提取领域不变特征;其次,实施增强的压缩操作以完整保留原图的结构与特征信息;最后,通过谱域增强对比学习,将因果不变特征注入压缩图中,确保压缩图保留原始图的因果信息。在五个公开数据集及新构建的FinReport数据集上的实验表明,TGCC在跨任务、跨域复杂场景下相比现有方法最高提升13.41%,并在6个数据集中的5个达到最优性能。

原文摘要 · Abstract (English)

The increasing scale of graph datasets has significantly improved the performance of graph representation learning methods, but it has also introduced substantial training challenges. Graph dataset condensation techniques have emerged to compress large datasets into smaller yet information-rich datasets, while maintaining similar test performance. However, these methods strictly require downstream applications to match the original dataset and task, which often fails in cross-task and cross-domain scenarios. To address these challenges, we propose a novel causal-invariance-based and transferable graph dataset condensation method, named TGCC, providing effective and transferable condensed datasets. Specifically, to preserve domain-invariant knowledge, we first extract domain causal-invariant features from the spatial domain of the graph using causal interventions. Then, to fully capture the structural and feature information of the original graph, we perform enhanced condensation operations. Finally, through spectral-domain enhanced contrastive learning, we inject the causal-invariant features into the condensed graph, ensuring that the compressed graph retains the causal information of the original graph. Experimental results on five public datasets and our novel FinReport dataset demonstrate that TGCC achieves up to a 13.41% improvement in cross-task and cross-domain complex scenarios compared to existing methods, and achieves state-of-the-art performance on 5 out of 6 datasets in the single dataset and task scenario.

图数据压缩因果学习可迁移性图神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。