构建多表问答评测基准,检验大模型在真实数据中的复杂推理能力
TQA-Bench: Evaluating LLMs for Multi-Table Question Answering
- 基于真实数据构建可变上下文长度的多表问答基准
- 覆盖20亿至6710亿参数模型,验证大模型在长文本中的推理表现
- 支持超越检索与模式匹配的符号化推理评估,适合数据密集型研究者
大型语言模型(LLMs)在复杂多模态数据管理任务中展现出巨大潜力,尤其在跨多表关系数据的问答(QA)方面。然而,由于关系数据结构分析的内在复杂性及序列化表格数据的潜在规模,系统评估LLMs在多表问答上的表现仍面临挑战。现有基准主要聚焦单表问答,难以捕捉金融、医疗、电商等实际场景中多表间的关联复杂性。本文提出TQA-Bench,一个基于真实公开数据集的长上下文分析型多表问答基准,支持8K至64K token的上下文长度采样,并引入符号扩展以评估超越检索与模式匹配的推理能力。我们系统评估了从20亿到6710亿参数的一系列LLMs。大量实验揭示了大模型在多表问答中的关键性能洞见,凸显其在复杂数据驱动环境中的挑战与机遇。
原文摘要 · Abstract (English)
The advance of large language models (LLMs) has unlocked great opportunities in complex multi-modal data management tasks, particularly in question answering (QA) over complicated multi-table relational data. Despite significant progress, systematically evaluating LLMs on multi-table QA remains a critical challenge due to the inherent complexity of analyzing the modality of relational data structures and the potentially large scale of serialized tabular data. Existing benchmarks primarily focus on single-table QA, failing to capture the intricacies of connections across multiple relational tables, as required in real-world domains such as finance, healthcare, and e-commerce. We present TQA-Bench, a long-context analytical multi-table QA benchmark derived from real-world public datasets, with a flexible sampling mechanism that varies context length (8K--64K tokens) and symbolic extensions for assessing reasoning beyond retrieval and pattern matching. We systematically evaluate a set of LLMs spanning model scales from 2 billion to 671 billion parameters. Our extensive experiments reveal critical insights into the performance of LLMs in multi-table QA, highlighting both challenges and opportunities for advancing their application in complex, data-driven environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。