arXiv:2601.04855cs.LGcs.AI2026-01被引 1

针对图神经网络缺失特征问题,提出更真实的数据与评估方法,并设计高效基线模型。

Rethinking GNNs and Missing Features: Challenges, Evaluation and a Robust Solution

  • 构建包含稠密语义特征的真实与合成数据集,突破传统稀疏数据局限
  • 设计非随机缺失机制评估协议,更贴近医疗等真实场景
  • 提出GNNmim基线模型,在多种缺失情形下表现稳定且优于专用架构

在医疗、传感器网络等现实场景中,处理缺失节点特征是图神经网络部署的关键挑战。现有研究多集中于相对温和的场景:高维但稀疏的节点特征,以及完全随机缺失(MCAR)机制下的不完整数据。我们理论证明,高稀疏性会显著限制缺失带来的信息损失,使所有模型看似鲁棒,难以有效比较性能。为此,我们引入一个合成数据集和三个真实世界数据集,其特征密集且具有语义意义。同时,超越MCAR假设,设计更符合实际的缺失机制评估方案,并提供缺失过程的显式理论假设及对不同方法的影响分析。基于此,我们提出GNNmim——一种简单但高效的节点分类基线方法。实验表明,GNNmim在多种数据集和缺失模式下均表现良好,与专门设计的模型相当甚至更优。

原文摘要 · Abstract (English)

Handling missing node features is a key challenge for deploying Graph Neural Networks (GNNs) in real-world domains such as healthcare and sensor networks. Existing studies mostly address relatively benign scenarios, namely benchmark datasets with (a) high-dimensional but sparse node features and (b) incomplete data generated under Missing Completely At Random (MCAR) mechanisms. For (a), we theoretically prove that high sparsity substantially limits the information loss caused by missingness, making all models appear robust and preventing a meaningful comparison of their performance. To overcome this limitation, we introduce one synthetic and three real-world datasets with dense, semantically meaningful features. For (b), we move beyond MCAR and design evaluation protocols with more realistic missingness mechanisms. Moreover, we provide a theoretical background to state explicit assumptions on the missingness process and analyze their implications for different methods. Building on this analysis, we propose GNNmim, a simple yet effective baseline for node classification with incomplete feature data. Experiments show that GNNmim is competitive with respect to specialized architectures across diverse datasets and missingness regimes.

图神经网络缺失数据稳健性评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。