arXiv:2511.09665cs.LG2025-11被引 4

仅用一张真实表格的自监督预训练,就能让表格模型实现跨领域泛化。

Generalization Can Emerge in Tabular Foundation Models From a Single Table

  • 在单张真实表格上进行自监督预训练,无需大规模数据集。
  • 仅用少量数据即可在多个异构基准上实现强迁移性能。
  • 构建的任务数量与质量是决定模型泛化能力的关键。

深度表格建模越来越多地依赖于上下文学习:推理时,模型接收一组(x,y)对作为上下文,预测新输入的标签而不更新权重。我们挑战了当前观点——广泛泛化需依赖大规模合成语料(如TabPFN先验)或真实数据集合(如TabDPT训练数据集),发现少量数据足以实现泛化。系统性地在多个多样化数据集上进行预训练和评估后,我们发现仅通过一张真实表格的自监督预训练,即可在异构基准间产生令人惊讶的强迁移能力。我们分析了数据中影响表格基础模型(TFM)跨域泛化的关键因素,并指出多数TFM共有的预训练流程中,从数据集中可构造的任务数量与质量是下游性能的核心。实验表明,仅需单个真实表,便可构建具备强大泛化能力的表格基础模型。

原文摘要 · Abstract (English)

Deep tabular modelling increasingly relies on in-context learning where, during inference, a model receives a set of $(x,y)$ pairs as context and predicts labels for new inputs without weight updates. We challenge the prevailing view that broad generalization here requires pre-training on large synthetic corpora (e.g., TabPFN priors) or a large collection of real data (e.g., TabDPT training datasets), discovering that a relatively small amount of data suffices for generalization. We find that simple self-supervised pre-training on just a \emph{single} real table can produce surprisingly strong transfer across heterogeneous benchmarks. By systematically pre-training and evaluating on many diverse datasets, we analyze what aspects of the data are most important for building a Tabular Foundation Model (TFM) generalizing across domains. We then connect this to the pre-training procedure shared by most TFMs and show that the number and quality of \emph{tasks} one can construct from a dataset is key to downstream performance.

表格模型自监督泛化能力零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。