构建真实工业场景的表格问答基准,推动大模型表推理能力提升
ReasonTabQA: A Comprehensive Benchmark for Table Question Answering from Real World Industrial Scenarios
- 构建跨30个工业领域的1932张真实表格数据集
- 引入可验证奖励机制的强化学习方法,显著提升推理准确率
- 支持带思维链与无思维链两种模式,适配工业级复杂场景
大型语言模型(LLM)的进展极大推动了基于表格的问答(TableQA)发展。然而,现有TableQA基准常忽略工业场景的复杂性,如多表结构、嵌套表头和超大规模数据,这些特性要求深层结构化推理,而当前方法尚未充分应对。为此,我们提出ReasonTabQA,一个涵盖30个工业领域(如能源、汽车)共1932张表格的大规模双语基准,提供高质量的最终答案与显式推理链标注,支持思考与非思考范式。此外,我们提出TabCodeRL,一种利用表格感知可验证奖励引导逻辑推理路径生成的强化学习方法。在ReasonTabQA及4个TableQA数据集上的大量实验表明,尽管TabCodeRL在开源LLM上取得显著性能提升,但在ReasonTabQA上仍存在明显性能差距,凸显真实工业场景TableQA的内在复杂性。
原文摘要 · Abstract (English)
Recent advancements in Large Language Models (LLMs) have significantly catalyzed table-based question answering (TableQA). However, existing TableQA benchmarks often overlook the intricacies of industrial scenarios, which are characterized by multi-table structures, nested headers, and massive scales. These environments demand robust table reasoning through deep structured inference, presenting a significant challenge that remains inadequately addressed by current methodologies. To bridge this gap, we present ReasonTabQA, a large-scale bilingual benchmark encompassing 1,932 tables across 30 industry domains such as energy and automotive. ReasonTabQA provides high-quality annotations for both final answers and explicit reasoning chains, supporting both thinking and no-thinking paradigms. Furthermore, we introduce TabCodeRL, a reinforcement learning method that leverages table-aware verifiable rewards to guide the generation of logical reasoning paths. Extensive experiments on ReasonTabQA and 4 TableQA datasets demonstrate that while TabCodeRL yields substantial performance gains on open-source LLMs, the persistent performance gap on ReasonTabQA underscores the inherent complexity of real-world industrial TableQA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。