arXiv:2604.22676cs.LG2026-04

用可解释的信号子空间分析图数据集特征,揭示模型决策背后的真正影响因素。

Operational Feature Fingerprints of Graph Datasets via a White-Box Signal-Subspace Probe

论文配图:Operational Feature Fingerprints of Graph Datasets via a White-Box Signal-Subspace Probe
图 1 · 摘自论文原文
  • 设计白盒探针WG-SRC,用固定信号字典替代神经网络消息传递
  • 在6个数据集上表现接近基线,且能分解出低通、高通等信号成分
  • 适合研究模型为何分类、如何优化图学习机制的学者使用

图神经网络在节点分类上表现优异,但其学习的消息传递过程会将节点属性、邻域平滑、图结构高频差异、类别几何和分类边界效应混杂于黑箱表示中,导致难以理解分类依据及数据集所需的学习机制。本文提出白盒信号子空间探针WG-SRC,以固定命名的图信号字典(包含原始特征、行归一化与对称归一化低通传播、高通图差异)替代学习的消息传递。结合Fisher坐标选择、类内PCA子空间、闭式多α岭分类与验证得分融合,使预测与分析基于显式类别子空间、能量可控维度与闭式线性决策。通过预测性能验证诊断结果。在六个节点分类数据集上,其表现与复现基线相当,并在对齐划分下获得正向平均增益。其生成的图谱将行为分解为原始特征、低通、高通、类别几何与岭边界成分,区分出以低通主导的Amazon图、混合高通与类别几何复杂的Chameleon行为,以及依赖原始特征或边界的WebKB图。对齐干预表明高通块可作为可移除噪声,图衍生信号应保留,岭修正在特定条件下重要。因此,WG-SRC既是可运行的白盒分类器,也是数据集指纹探测器,支持基于信号特征条件的黑箱组件行为分析。

原文摘要 · Abstract (English)

Graph neural networks achieve strong node-classification accuracy, but learned message passing entangles ego attributes, neighborhood smoothing, high-pass graph differences, class geometry, and classifier-boundary effects inside opaque representations. This obscures why nodes are classified as they are and which graph-learning mechanisms a dataset requires. We propose WG-SRC, a white-box signal-subspace probe for prediction and graph dataset diagnosis. WG-SRC replaces learned message passing with a fixed, named graph-signal dictionary containing raw features, row- and symmetric-normalized low-pass propagation, and high-pass graph differences. It combines Fisher coordinate selection, class-wise PCA subspaces, closed-form multi-alpha ridge classification, and validation-based score fusion, so prediction and analysis rely on explicit class subspaces, energy-controlled dimensions, and closed-form linear decisions. As a white-box graph-learning instrument, WG-SRC uses predictive performance to validate its diagnostics. Across six node-classification datasets, it remains competitive with reproduced baselines and achieves positive average gain under aligned splits. Its atlas decomposes behavior into raw-feature, low-pass, high-pass, class-geometric, and ridge-boundary components. The resulting fingerprints distinguish low-pass-dominated Amazon graphs, mixed high-pass and class-geometrically complex Chameleon behavior, and raw- or boundary-sensitive WebKB graphs. Aligned interventions show when high-pass blocks act as removable noise, when raw or graph-derived signals should be preserved, and when ridge correction matters. WG-SRC therefore serves both as a functioning white-box classifier and as a dataset-fingerprinting probe, enabling fingerprint-conditioned analysis of how black-box model components behave under different graph-signal conditions.

图神经网络可解释性信号分析白盒探针

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。