系统评估图信号对表格学习的增益,找出真正可靠的信号类型。
A Systematic Evaluation Protocol of Graph-Derived Signals for Tabular Machine Learning
- 构建可复现的评估协议,统一测试各类图信号效果。
- 在加密货币欺诈数据集上验证,识别出稳定提升性能的信号类别。
- 适合关注图神经网络在表格数据中应用的科研与工程人员。
尽管图衍生信号广泛应用于表格学习,现有研究多依赖有限实验设置和平均性能比较,导致观察到的提升在统计可靠性和鲁棒性方面缺乏充分探索。本文提出一种基于分类体系的实证分析方法,设计统一且可复现的评估协议,系统评估哪些类别的图衍生信号能带来统计显著且稳健的性能提升。该协议支持将多种图信号可控集成到表格学习流程中,包含自动化超参数优化、多种子统计评估、正式显著性检验及图结构扰动下的鲁棒性分析。通过在大规模、不平衡的加密货币欺诈检测数据集上的广泛案例研究,结果识别出持续有效提升性能的信号类别,并揭示了哪些图信号能反映欺诈性的结构模式。此外,鲁棒性分析显示不同信号对缺失或损坏关系数据的处理能力差异显著。这些发现对欺诈检测具有实用价值,也展示了该分类驱动评估协议在其他领域的可扩展性。
原文摘要 · Abstract (English)
While graph-derived signals are widely used in tabular learning, existing studies typically rely on limited experimental setups and average performance comparisons, leaving the statistical reliability and robustness of observed gains largely unexplored. Consequently, it remains unclear which signals provide consistent and robust improvements. This paper presents a taxonomy-driven empirical analysis of graph-derived signals for tabular machine learning. We propose a unified and reproducible evaluation protocol to systematically assess which categories of graph-derived signals yield statistically significant and robust performance improvements. The protocol provides an extensible setup for the controlled integration of diverse graph-derived signals into tabular learning pipelines. To ensure a fair and rigorous comparison, it incorporates automated hyperparameter optimization, multi-seed statistical evaluation, formal significance testing, and robustness analysis under graph perturbations. We demonstrate the protocol through an extensive case study on a large-scale, imbalanced cryptocurrency fraud detection dataset. The analysis identifies signal categories providing consistently reliable performance gains and offers interpretable insights into which graph-derived signals indicate fraud-discriminative structural patterns. Furthermore, robustness analyses reveal pronounced differences in how various signals handle missing or corrupted relational data. These findings demonstrate practical utility for fraud detection and illustrate how the proposed taxonomy-driven evaluation protocol can be applied in other application domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。