arXiv:2604.21696cs.LGcs.DB2026-04被引 3

构建表格嵌入通用评测基准,指导实际应用中模型选择。

Towards Universal Tabular Embeddings: A Benchmark Across Data Tasks

论文配图:Towards Universal Tabular Embeddings: A Benchmark Across Data Tasks
图 1 · 摘自论文原文
  • 提出TEmBed基准,覆盖单元格、行、列、表四级表示评估。
  • 不同任务和表示层级下,最佳模型表现差异显著。
  • 为真实场景中的表格嵌入选型提供实证依据。

表格基础模型旨在学习跨任务与领域的通用表格表示,支持表格检索、语义搜索及基于表格的预测等应用。尽管此类模型日益增多,但现有方法常在特定任务下评估,难以直接比较。为此,我们提出TEmBed(表格嵌入测试床),一个涵盖单元格、行、列、表四个表示层级的综合性评测基准。评估多种表格表示学习模型后发现,模型选择需根据具体任务与表示层级而定。研究结果为实际应用中嵌入模型的选择提供实用指导,并为开发更通用的表格表示模型奠定基础。

原文摘要 · Abstract (English)

Tabular foundation models aim to learn universal representations of tabular data that transfer across tasks and domains, enabling applications such as table retrieval, semantic search and table-based prediction. Despite the growing number of such models, it remains unclear which approach works best in practice, as existing methods are often evaluated under task-specific settings that make direct comparison difficult. To address this, we introduce TEmBed, the Tabular Embedding Test Bed, a comprehensive benchmark for systematically evaluating tabular embeddings across four representation levels: cell, row, column, and table. Evaluating a diverse set of tabular representation learning models, we show that which model to use depends on the task and representation level. Our results offer practical guidance for selecting tabular embeddings in real-world applications and lay the groundwork for developing more general-purpose tabular representation models.

表格嵌入基准评测表示学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。