系统梳理大模型时代表格问答的研究脉络与挑战。
Table Question Answering in the Era of Large Language Models: A Comprehensive Survey of Tasks, Methods, and Evaluation
- 按任务设置与挑战分类现有方法,梳理主流范式。
- 归纳大模型在表格问答中的应用策略与优劣。
- 揭示未被充分研究的前沿方向,适合研究者参考。
表格问答(TQA)旨在回答关于表格数据的自然语言问题,常伴随文本段落等附加上下文。该任务涵盖多种场景,涉及表格表示、问题/答案复杂度、模态和领域差异。尽管大语言模型(LLMs)带来了显著进展,但该领域仍缺乏对任务设定、核心挑战和方法趋势的系统性梳理,尤其在强化学习等新兴方向上。本综述聚焦基于大模型的TQA研究,系统化地分类现有基准与任务设置,根据应对挑战对建模策略进行分组分析,并评估其优缺点。同时,指出尚未被充分覆盖但具有时效性的研究空白。通过整合分散的研究线索并识别开放问题,本综述为TQA社区提供统一基础,促进对前沿技术的深入理解,并指导未来发展方向。
原文摘要 · Abstract (English)
Table Question Answering (TQA) aims to answer natural language questions about tabular data, often accompanied by additional contexts such as text passages. The task spans diverse settings, varying in table representation, question/answer complexity, modality involved, and domain. While recent advances in large language models (LLMs) have led to substantial progress in TQA, the field still lacks a systematic organization and understanding of task formulations, core challenges, and methodological trends, particularly in light of emerging research directions such as reinforcement learning. This survey addresses this gap by providing a comprehensive and structured overview of TQA research with a focus on LLM-based methods. We provide a comprehensive categorization of existing benchmarks and task setups. We group current modeling strategies according to the challenges they target, and analyze their strengths and limitations. Furthermore, we highlight underexplored but timely topics that have not been systematically covered in prior research. By unifying disparate research threads and identifying open problems, our survey offers a consolidated foundation for the TQA community, enabling a deeper understanding of the state of the art and guiding future developments in this rapidly evolving area.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。