提出GRL-Safety评估图表示学习在多种部署压力下的安全性
On the Safety of Graph Representation Learning

- 构建多维度安全评测框架,覆盖五类关键风险
- 12种方法在25个数据集上测试,发现基础模型无全面优势
- 揭示实际部署中仍存挑战,需新训练目标或适应机制
图表示学习(GRL)已从仅依赖拓扑的嵌入发展到任务特定的监督GNN,再到可复用表示和图基础模型(GFMs)。然而现有评估主要关注干净环境下的迁移、适配和任务覆盖,尚未明确当部署压力影响图信号、上下文、标签支持、结构分组或预测证据时,GRL方法是否仍可靠。本文提出GRL-Safety,一个针对GRL的多轴安全评估基准。该基准在标准化条件下对12种代表性方法(涵盖拓扑嵌入、监督GNN、自监督图模型及GFMs)在25个图数据集上进行评估,保持方法原生适配性。评估涵盖五个安全维度:噪声鲁棒性、分布外泛化、类别不平衡、公平性与可解释性,并提供各轴及子条件下的细粒度报告,而非单一综合分数。分析得出三个跨轴洞见:第一,安全表现由表示设计与受压因素的交互决定,而非方法族本身;第二,基础时代方法仅在特定轴上表现突出,无全局优势;第三,即使最佳方法在某些部署场景仍难以应对,暴露出需新鲁棒性、适配或训练目标的能力缺口。基准、评估协议与代码已开源:https://github.com/GXG-CS/GRL-Safety。
原文摘要 · Abstract (English)
Graph representation learning (GRL) has evolved from topology-only graph embeddings to task-specific supervised GNNs, and more recently to reusable representations and graph foundation models (GFMs). However, existing evaluations mainly measure clean transfer, adaptation, and task coverage. It remains unclear whether GRL methods stay reliable when deployment stresses affect graph signals, graph contexts, label support, structural groups, or predictive evidence. We introduce GRL-Safety, a multi-axis safety evaluation benchmark for GRL. GRL-Safety evaluates twelve representative methods, spanning topology-only embedding methods, supervised GNNs, self-supervised graph models, and GFMs, on twenty-five graph datasets under standardized evaluation conditions while preserving method-native adaptation. The evaluation covers five safety axes: corruption robustness, OOD generalization, class imbalance, fairness, and interpretation, with per-axis and sub-condition reporting rather than a single aggregate score. Our analysis yields three cross-axis insights that can inspire future research. First, safety behavior is shaped by the interaction between representation design and the stressed graph factor, rather than by method family alone. Second, foundation-era methods show axis-specific strengths rather than broad safety dominance. Third, several deployment regimes remain difficult even for the best evaluated method, revealing capability gaps that require new robustness, adaptation, or training objectives beyond model selection. The benchmark, evaluation protocols, and code are available at: https://github.com/GXG-CS/GRL-Safety.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。