arXiv:2607.25273cs.LG2026-07中稿 · MAPR 2026

针对异质图设计更精准的不确定性预测,提升GNN可靠性

HeAD-CP: Heterophily-Aware Diffused Conformal Prediction Sets for Graph Neural Networks

  • 基于GNN输出自适应计算局部同质性,动态调整扩散系数
  • 在异质图上平均预测集大小比DAPS减少10.3%,且覆盖保证不变
  • 特别适合异质图场景,对模型不确定性评估有显著提升

置信度预测(CP)提供无需分布假设的不确定性量化,其在图数据上的扩展是当前研究热点。现有的图感知扩散基线DAPS通过统一系数λ沿边传播自适应预测集的非符合度分数,但该方法隐含假设图具有同质性,在异质图上表现不佳,使平均预测集大小相对基础APS增大高达10.6%。为此,本文提出HeAD-CP,一种基于节点的扩散变体族,其扩散系数由GNN softmax导出的无标签局部同质性估计决定。三个变体——符号γ、边兼容性以及修正版DAPS——分别在极端异质、中等异质和中高同质场景下最优,且均保持边际覆盖保证。在十个基准数据集上,HeAD-CP家族在所有数据集上均不劣于或优于基础APS,而DAPS仅在六个数据集上胜出。对家族进行事后优化可显著超越DAPS(8/10数据集,p<0.01),尤其在异质图上提升明显(如Texas达10.3%);在两个同质图(CiteSeer, PubMed)上虽仍略优,但优势极小(≤0.002),统计不显著(CiteSeer: p=0.23)。如何设计一个接近此最优选择器的校准无标签筛选器,是当前主要的实证挑战。

原文摘要 · Abstract (English)

Conformal prediction (CP) provides distribution-free uncertainty quantification, and its extension to graphs is an active research direction. Diffused Adaptive Prediction Sets (DAPS) is a widely used graph-aware diffusion baseline, propagating Adaptive Prediction Sets (APS) non-conformity scores along edges with a uniform coefficient $λ$. We identify a fundamental shortcoming of this design: the uniform low-pass diffusion presupposes graph homophily and proves detrimental on heterophilic graphs, enlarging the mean prediction-set size by up to 10.6% relative to plain APS. To mitigate this, we propose HeAD-CP, a family of node-wise diffusion variants whose coefficients are determined by a label-free local-homophily estimate derived from the GNN softmax. Three variants, namely signed-$γ$, edge-compatibility, and a DAPS-baseline-with-correction, are most effective at extreme heterophily, intermediate heterophily, and moderate-to-high homophily, respectively, and all preserve the marginal coverage guarantee. On ten benchmarks, the HeAD-CP family stays at or below plain APS on every dataset, while DAPS exceeds APS on six. The post-hoc oracle over the family improves over DAPS on 8/10 datasets at $p<0.01$ (paired Wilcoxon), with the largest gains on heterophilic graphs (10.3% on Texas); on the two homophilic datasets where DAPS still wins (CiteSeer, PubMed), it retains a marginal advantage of at most 0.002, statistically insignificant on CiteSeer ($p=0.23$). Designing a calibrated label-free selector that approaches this oracle is the main outstanding empirical question.

图神经网络不确定性量化异质图置信度预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。