首个统一视觉表格多模态基准,推动医疗等高风险领域研究
VT-Bench: A Unified Benchmark for Visual-Tabular Multi-Modal Learning

- 构建跨9大领域的14个数据集,覆盖超75.6万样本
- 评估23种模型,揭示视觉表格学习的显著挑战
- 适合关注医疗、工业等场景的多模态研究者
多模态学习在视觉-文本任务中备受关注,但视觉-表格数据在医疗、工业等高风险领域中仍缺乏系统研究。本文提出VT-Bench,首个标准化视觉-表格判别预测与生成推理任务的统一基准。该基准整合了来自9个领域的14个数据集,涵盖医疗、宠物、媒体、交通等,总样本量超过75.6万。我们评估了23种代表性模型,包括单模态专家、专用视觉-表格模型、通用视觉-语言模型(VLMs)以及工具增强方法,揭示了视觉-表格学习中的显著挑战。我们相信,VT-Bench将推动社区构建更强大的多模态视觉-表格基础模型。
原文摘要 · Abstract (English)
Multi-model learning has attracted great attention in visual-text tasks. However, visual-tabular data, which plays a pivotal role in high-stakes domains like healthcare and industry, remains underexplored. In this paper, we introduce \textit{VT-Bench}, the first unified benchmark for standardizing vision-tabular discriminative prediction and generative reasoning tasks. VT-Bench aggregates 14 datasets across 9 domains (medical-centric, while covering pets, media, and transportation) with over 756K samples. We evaluate 23 representative models, including unimodal experts, specialized visual-tabular models, general-purpose vision-language models (VLMs), and tool-augmented methods, highlighting substantial challenges of visual-tabular learning. We believe VT-Bench will stimulate the community to build more powerful multi-modal vision-tabular foundation models. Benchmark: https://github.com/Ziyi-Jia990/VT-Bench
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。