为表格数据中的文本特征设计新评测基准,推动模型能力评估
Towards Benchmarking Foundation Models for Tabular Data With Text
- 提出简单有效的文本融合策略,适配传统表格处理流程
- 人工筛选真实世界含语义文本的表格数据集用于评测
- 填补现有基准对文本特征支持的空白,适合研究表格多模态模型者
表格数据的基础模型快速发展,越来越多研究尝试扩展其支持自由文本特征。然而,现有表格数据基准极少包含文本列,且在真实场景中寻找富含语义的文本特征表格数据颇具挑战。本文提出一系列简单而有效的消融式策略,将文本信息融入传统表格处理流程。同时,通过手动整理一批具有实际意义文本特征的真实世界表格数据集,评测当前最先进的表格基础模型在处理文本数据时的表现。本研究是提升表格数据基础模型文本能力评测的重要一步。
原文摘要 · Abstract (English)
Foundation models for tabular data are rapidly evolving, with increasing interest in extending them to support additional modalities such as free-text features. However, existing benchmarks for tabular data rarely include textual columns, and identifying real-world tabular datasets with semantically rich text features is non-trivial. We propose a series of simple yet effective ablation-style strategies for incorporating text into conventional tabular pipelines. Moreover, we benchmark how state-of-the-art tabular foundation models can handle textual data by manually curating a collection of real-world tabular datasets with meaningful textual features. Our study is an important step towards improving benchmarking of foundation models for tabular data with text.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。