构建多表推理的洞察生成基准与评估框架,解决大模型跨表分析难题。
MT-RAIG: Novel Benchmark and Evaluation Framework for Retrieval-Augmented Insight Generation over Multiple Tables
- 提出MT-RAIG Bench,评测系统从多张未知表格中生成洞察的能力。
- 设计细粒度评估框架,使机器评价更贴近人类对分析质量的判断。
- 发现顶尖大模型在复杂多表推理上仍表现不佳,适合研究者挑战新方法。
表式推理近年已从事实问答拓展至需要合成隐含知识的洞察生成任务,要求系统提供可解释的分析。然而现有研究仅限于单张已知表格的场景,无法应对用户从多张未知表格中获取综合洞察的需求。为此,我们提出MT-RAIG Bench,用于评估多表检索增强型洞察生成能力。同时,为克服现有自动评估方法在表格领域表现不佳的问题,进一步引入细粒度评估框架MT-RAIG Eval,其评价结果与人工质量判断更一致。通过大量实验发现,即便前沿大模型在复杂多表推理任务中仍存在显著困难,确立了该基准作为未来研究的挑战性测试平台。
原文摘要 · Abstract (English)
Recent advancements in table-based reasoning have expanded beyond factoid-level QA to address insight-level tasks, where systems should synthesize implicit knowledge in the table to provide explainable analyses. Although effective, existing studies remain confined to scenarios where a single gold table is given alongside the user query, failing to address cases where users seek comprehensive insights from multiple unknown tables. To bridge these gaps, we propose MT-RAIG Bench, design to evaluate systems on Retrieval-Augmented Insight Generation over Mulitple-Tables. Additionally, to tackle the suboptimality of existing automatic evaluation methods in the table domain, we further introduce a fine-grained evaluation framework MT-RAIG Eval, which achieves better alignment with human quality judgments on the generated insights. We conduct extensive experiments and reveal that even frontier LLMs still struggle with complex multi-table reasoning, establishing our MT-RAIG Bench as a challenging testbed for future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。