arXiv:2410.13516cs.LG2024-10中稿 · NeurIPS被引 13

无需清洗数据,端到端预训练表格模型,提升可扩展性。

PORTAL: Scalable Tabular Foundation Models via Content-Specific Tokenization

  • 按行独立预训练,支持多源异构表格数据
  • 在未清洗的在线数据上训练,性能媲美顶尖方法
  • 适合大规模真实场景表格任务,无需结构化要求

表格数据的自监督学习试图将自然语言和图像领域的进展应用于多样化的表格领域。然而,现有技术常难以整合跨领域数据,且需要数据清洗或特定结构要求,限制了预训练数据集的可扩展性。我们提出PORTAL(Pretraining One-Row-at-a-Time for All tabLes),一种无需清洗或预处理即可处理多种数据模态的框架。该简单而强大的方法可在在线收集的数据集上有效预训练,并在复杂分类与回归任务上微调至达到当前最优水平。本工作为大规模表格数据的自监督学习提供了实际进步。

原文摘要 · Abstract (English)

Self-supervised learning on tabular data seeks to apply advances from natural language and image domains to the diverse domain of tables. However, current techniques often struggle with integrating multi-domain data and require data cleaning or specific structural requirements, limiting the scalability of pre-training datasets. We introduce PORTAL (Pretraining One-Row-at-a-Time for All tabLes), a framework that handles various data modalities without the need for cleaning or preprocessing. This simple yet powerful approach can be effectively pre-trained on online-collected datasets and fine-tuned to match state-of-the-art methods on complex classification and regression tasks. This work offers a practical advancement in self-supervised learning for large-scale tabular data.

表格预训练自监督学习无清洗数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。