arXiv:2511.15941cs.LGcs.AI2025-11KDD被引 3

iLTM统一树模型与神经网络,提升表格数据学习性能。

iLTM: Integrated Large Tabular Model

  • 融合树嵌入与元训练超网络的统一架构
  • 在1800+数据集上预训练,小到大尺寸任务均表现优异
  • 少调参即可超越主流树模型和深度模型

表格数据支撑科学、工业和公共服务中的决策。尽管深度学习进展迅速,但其优势尚未充分惠及表格领域,梯度提升决策树(GBDTs)仍是实际首选。我们提出iLTM——一种集成大型表格模型,将树衍生嵌入、维度无关表示、元训练超网络、多层感知机(MLPs)及检索机制整合于单一架构中。该模型在超过1,800个异构分类数据集上预训练,可在从小样本到大规模高维任务的各类表格分类与回归任务中持续实现优越性能。经轻量微调后,元训练超网络可迁移至回归任务,表现匹配或超越强基准。大量实验表明,iLTM优于经过精心调优的GBDTs和领先的深度表格模型,且所需任务特异性调参更少。通过弥合树方法与神经方法之间的差距,iLTM为鲁棒、灵活且可扩展的表格基础模型提供了新框架。

原文摘要 · Abstract (English)

Tabular data underpins decisions across science, industry, and public services. Despite rapid progress, advances in deep learning have not fully carried over to the tabular domain, where gradient-boosted decision trees (GBDTs) remain a default choice in practice. We present iLTM, an integrated Large Tabular Model that unifies tree-derived embeddings, dimensionality-agnostic representations, a meta-trained hypernetwork, multilayer perceptrons (MLPs), and retrieval within a single architecture. Pre-trained on more than 1,800 heterogeneous classification datasets, iLTM achieves consistently superior performance across tabular classification and regression tasks, from small datasets to large and high-dimensional tasks. After light fine-tuning, the meta-trained hypernetwork transfers to regression targets, matching or surpassing strong baselines. Extensive experiments show that iLTM outperforms well-tuned GBDTs and leading deep tabular models while requiring less task-specific tuning. By bridging the gap between tree-based and neural methods, iLTM offers a new framework for tabular foundation models for robust, adaptable, and scalable tabular learning.

表格模型深度学习元学习预训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。