arXiv:2607.12310cs.CLcs.AI2026-07中稿 · COLM

构建跨数据湖的问答基准,测试系统从杂乱数据中发现并整合信息的能力。

LakeQuest: A Three-Domain Benchmark for Grounded Question Answering across Data Lakes

论文配图:LakeQuest: A Three-Domain Benchmark for Grounded Question Answering across Data Lakes
图 1 · 摘自论文原文
  • 设计三领域真实数据湖的问答对,模拟复杂发现过程
  • 9846个问题验证端到端检索与合成性能,暴露关键失败模式
  • 适合研究智能体问答、多模态知识融合的开发者与学者

现代问答系统在结构清晰的数据上表现优异,但现实中的企业与科学数据常以异构、弱结构的形式存在,包含表格、文本段落和关联元数据。现有基准忽略这一噪声发现过程,无法评估端到端性能。为此,我们提出LakeQuest,一个经人工验证的9846个问答对的基准,用于评估真实数据湖上的端到端检索-合成流程。该基准涵盖三个不同领域(AI/ML元数据、零售银行、多模态生物医学药物信息),每个问题均配有精确的、模态感知的证据指针。通过分离源发现与跨模态合成,LakeQuest揭示了当前QA系统在元数据图关系链、银行账本政策定位以及生物医学场景联合表格问答中的普遍缺陷,表明未来智能体问答需更强的发现机制与忠实的跨文件组合能力。

原文摘要 · Abstract (English)

While modern question answering (QA) systems excel on clean, schema-aligned corpora, real-world knowledge is rarely so neatly packaged. Answering questions over enterprise and scientific data lakes requires systems to navigate heterogeneous, weakly structured collections of tables, passages, and linked metadata. Current benchmarks abstract away this noisy discovery process, failing to evaluate end-to-end performance. To bridge this gap, we introduce LakeQuest, a human-validated benchmark of 9,846 QA pairs designed to evaluate the end-to-end retrieve-and-synthesize pipeline over realistic data lakes. LakeQuest spans three diverse domains (AI/ML metadata, retail banking, and multimodal biomedical drug information) and pairs every question with exact, modality-aware evidence pointers. By isolating source discovery from cross-modal synthesis, LakeQuest exposes critical failure modes in modern QA systems. Our baseline evaluations, including standard Retrieval-Augmented Generation (RAG) and agentic tool-use methods, reveal that high-quality retrieval does not guarantee correct reasoning. Systems consistently struggle with relation chaining in metadata graphs, policy grounding in bank ledgers, and joint tabular QA in biomedical contexts, highlighting the need for robust discovery and faithful cross-file composition mechanisms in future agentic QA systems.

问答系统数据湖智能体多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。