arXiv:2410.12034cs.LGcs.AI2024-10综述被引 57

综述深度学习在表格数据上的演进,涵盖模型、方法与应用挑战。

A Survey on Deep Tabular Learning

  • 从全连接网络到注意力机制,融合特征嵌入与混合架构应对表格数据复杂性。
  • 先进模型如TabNet、SAINT提升可解释性与计算效率,支持大规模数据处理。
  • 适合关注表格数据建模、可解释性与工业应用的研究者与工程师阅读。

表格数据广泛应用于医疗、金融和交通等行业,因其异质性和缺乏空间结构,给深度学习带来独特挑战。本综述回顾了深度学习模型在表格数据上的演进,从早期的全连接网络(FCNs)到先进的架构如TabNet、SAINT、TabTranSELU和MambaNet。这些模型引入注意力机制、特征嵌入和混合架构以应对表格数据的复杂性。TabNet采用序列注意力实现实例级特征选择,增强可解释性;SAINT结合自注意力与样本间注意力,捕捉特征与数据点间的复杂交互,同时提升可扩展性并降低计算开销。混合架构如TabTransformer和FT-Transformer将注意力机制与多层感知机(MLPs)结合,有效处理分类与数值数据,其中FT-Transformer适配变压器用于表格数据集。研究持续致力于在大规模数据上平衡性能与效率。图基模型如GNN4TDL和GANDALF结合神经网络与决策树或图结构,通过先进正则化技术增强特征表示,缓解小样本过拟合问题。扩散模型如表格式去噪扩散概率模型(TabDDPM)生成合成数据以解决数据稀缺,提升模型鲁棒性。类似地,TabPFN和Ptab利用预训练语言模型,融入迁移学习与自监督技术于表格任务。本综述总结关键进展,并指出未来在可扩展性、泛化性与可解释性方面的研究方向。

原文摘要 · Abstract (English)

Tabular data, widely used in industries like healthcare, finance, and transportation, presents unique challenges for deep learning due to its heterogeneous nature and lack of spatial structure. This survey reviews the evolution of deep learning models for tabular data, from early fully connected networks (FCNs) to advanced architectures like TabNet, SAINT, TabTranSELU, and MambaNet. These models incorporate attention mechanisms, feature embeddings, and hybrid architectures to address tabular data complexities. TabNet uses sequential attention for instance-wise feature selection, improving interpretability, while SAINT combines self-attention and intersample attention to capture complex interactions across features and data points, both advancing scalability and reducing computational overhead. Hybrid architectures such as TabTransformer and FT-Transformer integrate attention mechanisms with multi-layer perceptrons (MLPs) to handle categorical and numerical data, with FT-Transformer adapting transformers for tabular datasets. Research continues to balance performance and efficiency for large datasets. Graph-based models like GNN4TDL and GANDALF combine neural networks with decision trees or graph structures, enhancing feature representation and mitigating overfitting in small datasets through advanced regularization techniques. Diffusion-based models like the Tabular Denoising Diffusion Probabilistic Model (TabDDPM) generate synthetic data to address data scarcity, improving model robustness. Similarly, models like TabPFN and Ptab leverage pre-trained language models, incorporating transfer learning and self-supervised techniques into tabular tasks. This survey highlights key advancements and outlines future research directions on scalability, generalization, and interpretability in diverse tabular data applications.

表格数据深度学习可解释性模型综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。