构建多维度表格嵌入评测基准,揭示单一任务评估的局限性
TEmBed-T: A Multi-Dimensional Benchmark for Table-Level Embeddings

- 扩展TEmBed测试平台,覆盖多种下游任务
- 实验证明无模型在所有任务中表现最优
- 适合研究表格表示学习与评测方法的学者
表格数据是主要的结构化数据形式,学习表格表示已成为核心研究方向。表格级嵌入支撑了表格检索、数据湖发现和表格分类等多种应用。尽管其重要性突出,当前对不同嵌入方法在各类任务中的表现仍缺乏系统理解,因此需要开展系统性评估。本文通过扩展近期提出的表格嵌入测试平台TEmBed,将其从单一检索任务扩展至涵盖多个互补属性的综合评估体系。基于对TEmBed模型池的实证研究,结果表明:没有任何一个模型在所有任务上均表现最优,说明表格嵌入质量不能仅由检索性能衡量。
原文摘要 · Abstract (English)
Tabular data is the dominant structured-data modality, and learning table representations has become a core research direction. Table-level embeddings in particular underpin a wide range of applications, including table retrieval, data lake discovery, and table classification. Despite their importance, there is still limited understanding of how different embedding approaches behave across tasks, making systematic evaluation and analysis essential. In this work, we introduce a systematic evaluation of table-level embeddings that captures several complementary properties required for downstream effectiveness. We realize this evaluation by extending TEmBed, a recently proposed testbed for tabular embeddings, whose table-level coverage is currently limited to a single retrieval task. An empirical study over the TEmBed model pool confirms that no single model excels across all tasks, demonstrating that table-level embedding quality cannot be reduced to retrieval alone.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。