梳理表格理解的最新进展与挑战,揭示大模型在处理复杂表格时的瓶颈。
Tabular Data Understanding with LLMs: A Survey of Recent Advances and Challenges
- 构建表格输入表示与任务的分类体系,系统归纳研究方向。
- 指出当前多数任务依赖检索,缺乏深层推理能力。
- 适合关注大模型表格理解、多模态分析的研究者阅读。
表格因其复杂的二维结构和灵活的表现形式,在大语言模型(LLMs)和多模态大语言模型(MLLMs)中受到广泛关注。与线性文本不同,表格涵盖从结构化数据库表到复杂嵌套电子表格等多种格式,用途各异。这种多样化的格式与用途催生了专用方法与任务,而非通用方案,导致表格理解任务的导航变得困难。为应对这些挑战,本文通过构建表格输入表示的分类体系,并介绍表格理解任务,梳理关键概念。我们指出了领域内几个关键空白:(1) 以检索为主的任务占主导,模型推理能力仅限于数学与逻辑操作;(2) 模型在处理复杂表格结构、大规模表格、长上下文或跨表格场景时面临显著挑战;(3) 模型在不同表格表示与格式间的泛化能力有限。
原文摘要 · Abstract (English)
Tables have gained significant attention in large language models (LLMs) and multimodal large language models (MLLMs) due to their complex and flexible structure. Unlike linear text inputs, tables are two-dimensional, encompassing formats that range from well-structured database tables to complex, multi-layered spreadsheets, each with different purposes. This diversity in format and purpose has led to the development of specialized methods and tasks, instead of universal approaches, making navigation of table understanding tasks challenging. To address these challenges, this paper introduces key concepts through a taxonomy of tabular input representations and an introduction of table understanding tasks. We highlight several critical gaps in the field that indicate the need for further research: (1) the predominance of retrieval-focused tasks that require minimal reasoning beyond mathematical and logical operations; (2) significant challenges faced by models when processing complex table structures, large-scale tables, length context, or multi-table scenarios; and (3) the limited generalization of models across different tabular representations and formats.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。