对比大模型在科学与非科学表格上的理解能力,发现处理科学表格时表现显著下降。
Table Understanding and (Multimodal) LLMs: A Cross-Domain Case Study on Scientific vs. Non-Scientific Data
- 跨领域比较文本与多模态大模型对表格的理解效果
- 科学表格上模型准确率下降约30%,图像格式影响较小
- 适合关注表格理解瓶颈的研究者与实际应用开发者
表格是科研、商业、医疗和教育中表示结构化数据最常用工具之一。尽管大模型在下游任务中表现出色,但其处理表格数据的效率仍缺乏深入研究。本文通过跨领域、跨模态评估,考察了文本与多模态大模型在科学与非科学语境表格上的表现,并分析其在图像与文本形式表格中的鲁棒性。同时开展可解释性分析,衡量上下文使用与输入相关性。我们提出TableEval基准,包含来自学术论文、维基百科和财务报告的3017张表格,每张表格提供五种格式:图像、字典、HTML、XML和LaTeX。结果表明,虽然大模型在不同模态间保持一定鲁棒性,但在处理科学表格时面临显著挑战。
原文摘要 · Abstract (English)
Tables are among the most widely used tools for representing structured data in research, business, medicine, and education. Although LLMs demonstrate strong performance in downstream tasks, their efficiency in processing tabular data remains underexplored. In this paper, we investigate the effectiveness of both text-based and multimodal LLMs on table understanding tasks through a cross-domain and cross-modality evaluation. Specifically, we compare their performance on tables from scientific vs. non-scientific contexts and examine their robustness on tables represented as images vs. text. Additionally, we conduct an interpretability analysis to measure context usage and input relevance. We also introduce the TableEval benchmark, comprising 3017 tables from scholarly publications, Wikipedia, and financial reports, where each table is provided in five different formats: Image, Dictionary, HTML, XML, and LaTeX. Our findings indicate that while LLMs maintain robustness across table modalities, they face significant challenges when processing scientific tables.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。