构建百万级数据湖问答基准,考验模型搜索与推理能力。
LakeQA: An Exploratory QA Benchmark over a Million-Scale Data Lake

- 基于9.5TB异构数据构建搜索驱动的多跳问答任务
- 顶尖模型在该基准上仅达18.37%精确匹配率
- 适合评估真实场景下LLM数据检索与分析能力
近期大语言模型在有显式证据或可简单检索的阅读理解问答任务中进展迅速。然而,现实问题通常缺乏准确证据文档,有效信息隐藏于海量数据湖中,需先搜索后推理。目前缺乏同时考察搜索与推理能力的综合性基准。为此,我们提出LakeQA,一个面向数据湖的以搜索为核心的问答基准,强调搜索与推理的联合能力。数据源自约9.5TB的维基百科及开源政府数据,涵盖结构化与非结构化文本。每条任务由至少一位博士级专家标注,要求长时程多跳推理,需发现正确文档并整合跨源证据。在七种前沿大模型上的实验表明,该任务极具挑战性:例如,GPT-5.2在该基准上仅获18.37%精确匹配分数。总体而言,LakeQA为开发能从现代数据湖中查找并分析数据的LLM代理提供了真实测试平台。
原文摘要 · Abstract (English)
Recent large language models (LLMs) have shown rapid progress in reading-based question answering (QA), where evidence is explicitly provided or can be trivially retrieved. In contrast, real-world questions are often not paired with accurate evidence documents. The useful evidence resides in massive data lakes, making search a prerequisite for answering. However, there is a lack of comprehensive benchmarks that require both searching and reasoning over large data lakes. To this end, we introduce LakeQA, a comprehensive benchmark for search-centric question answering over data lakes that jointly emphasizes searching and reasoning capabilities. LakeQA is built on a heterogeneous collection of approximately 9.5 TB of text resources from Wikipedia and open-source government data, spanning structured and unstructured data. To ensure task quality, each sample is annotated by at least one Ph.D.-level expert. Each task requires long-horizon multi-hop reasoning with implicit intermediate steps: agents need to discover the correct documents and then compose evidence across sources to produce the answer. Experimental results on seven frontier LLMs demonstrate that LakeQA is challenging. For instance, GPT-5.2 achieves only an exact-match score of 18.37% on LakeQA. Overall, LakeQA provides a realistic testbed for developing LLM agents that can both find and analyze data in modern data lakes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。