arXiv:2510.04543cs.LGstat.ML2025-10中稿 · TMLR 2026

现有图模型难以捕捉真实特征关系,结构建模才是提升预测的关键。

The Role of Feature Interactions in Graph-based Tabular Deep Learning

  • 用已知结构的合成数据检验图模型对特征交互的建模能力
  • 现有方法边恢复率接近随机,无法识别有效特征关系
  • 强制使用真实结构后准确率显著提升,说明结构建模很重要

表格数据的精准预测依赖于捕捉复杂且数据集特异的特征交互。基于注意力机制和图神经网络的图基表格深度学习(GTDL)方法通过将特征交互建模为图来提升预测性能。本文分析了这些方法如何建模特征交互。当前的GTDL方法主要关注预测准确性,常忽视底层图结构的精确建模。利用具有已知真值图结构的合成数据集,我们发现现有GTDL方法无法恢复有意义的特征交互,其边恢复率接近随机水平。这表明当前使用的注意力机制和消息传递方案未能有效捕捉特征交互。进一步地,当引入真实的交互结构时,预测准确率显著提升。结果强调,GTDL方法应优先考虑图结构的精确建模,因为这直接带来更好的预测效果。

原文摘要 · Abstract (English)

Accurate predictions on tabular data rely on capturing complex, dataset-specific feature interactions. Attention-based methods and graph neural networks, referred to as graph-based tabular deep learning (GTDL), aim to improve predictions by modeling these interactions as a graph. In this work, we analyze how these methods model the feature interactions. Current GTDL approaches primarily focus on optimizing predictive accuracy, often neglecting the accurate modeling of the underlying graph structure. Using synthetic datasets with known ground-truth graph structures, we find that current GTDL methods fail to recover meaningful feature interactions, as their edge recovery is close to random. This suggests that the attention mechanism and message-passing schemes used in GTDL do not effectively capture feature interactions. Furthermore, when we impose the true interaction structure, we find that the predictive accuracy improves. This highlights the need for GTDL methods to prioritize accurate modeling of the graph structure, as it leads to better predictions.

图神经网络特征交互表格数据可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。