构建首个大规模图文表格学习基准,验证任务感知表征的普适优势。
MulTaBench: Benchmarking Multimodal Tabular Learning with Text and Image

- 设计40个互补模态数据集,聚焦真实场景下的联合建模需求。
- 实证表明任务感知表征在文本与图像模态上均显著提升性能。
- 适合研究多模态基础模型、表征学习及医疗电商等应用方向。
表格基础模型在监督式表格学习中已达到最新水平,通过预训练学习数值与类别结构化数据的通用表征。然而,它们缺乏对文本和图像等非结构化模态的原生支持,依赖冻结的预训练嵌入处理这些模态。在现有多模态表格学习基准上,我们发现微调嵌入可提升性能。但现有基准往往仅关注模态共现,导致数据集间方差大,掩盖了任务特定微调的优势。为此,我们提出MulTaBench,一个包含40个数据集的基准,图像-表格与文本-表格任务各20个。我们聚焦于模态提供互补预测信号且通用嵌入会丢失关键信息的任务,强调需任务感知表征。实验结果表明,任务感知表征微调的收益在文本与图像模态、多种表格学习器、编码器规模和嵌入维度下均具泛化性。MulTaBench是迄今最大的图像-表格基准,覆盖医疗、电商等高影响力领域。其旨在推动新型架构与联合建模方法研究,促进多模态表格基础模型的发展。
原文摘要 · Abstract (English)
Tabular Foundation Models have recently established the state of the art in supervised tabular learning, by leveraging pretraining to learn generalizable representations of numerical and categorical structured data. However, they lack native support for unstructured modalities such as text and image, and rely on frozen, pretrained embeddings to process them. On established Multimodal Tabular Learning benchmarks, we show that tuning the embeddings to the task improves performance. Existing benchmarks, however, often focus on the mere co-occurrence of modalities; this leads to high variance across datasets and masks the benefits of task-specific tuning. To address this gap, we introduce MulTaBench, a benchmark of 40 datasets, split equally between image-tabular and text-tabular tasks. We focus on predictive tasks where the modalities provide complementary predictive signal, and where generic embeddings lose critical information, necessitating Target-Aware Representations that are aligned with the task. Our experimental results demonstrate that the gains from target-aware representation tuning generalize across both text and image modalities, several tabular learners, encoder scales, and embedding dimensions. MulTaBench constitutes the largest image-tabular benchmarking effort to date, spanning high-impact domains such as healthcare and e-commerce. It is designed to enable the research of novel architectures which incorporate joint modeling and target-aware representations, paving the way for the development of novel Multimodal Tabular Foundation Models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。