构建可验证的深度研究问答评估体系,解决复杂问题回答的评测难题。
Total Recall QA: A Verifiable Evaluation Suite for Deep Research Agents
- 基于结构化知识库与文本语料构建精准可验证的问答任务
- 引入全召回问答新范式,实现大规模高质量数据生成
- 面向真实场景设计基准测试,适配研究型AI模型评估
深度研究代理作为基于大语言模型的系统,能够通过多步信息检索与推理,在开放域中综合多个信息源回答复杂问题。然而,由于任务复杂性高,现有评估方法仍面临根本性挑战。本文提出深度研究代理评估所需的核心要求与可选属性,发现现有基准均未满足全部要求。受TREC全召回赛道启发,我们提出全召回问答任务,并构建符合标准的评估框架:利用结构化知识库与文本语料生成单答案、全召回型问题,实现精确评估与相关性判断。基于此框架,我们构建了TRQA基准,其数据来源于Wikidata-Wikipedia真实来源,以及合成生成的电商知识库与语料,有效避免数据污染。我们在该基准上对代表性检索器与深度研究模型进行测评,建立了可复现的检索与端到端基线结果,为未来研究提供统一评估标准。
原文摘要 · Abstract (English)
Deep research agents have emerged as LLM-based systems designed to perform multi-step information seeking and reasoning over large, open-domain sources to answer complex questions by synthesizing information from multiple information sources. Given the complexity of the task and despite various recent efforts, evaluation of deep research agents remains fundamentally challenging. This paper identifies a list of requirements and optional properties for evaluating deep research agents. We observe that existing benchmarks do not satisfy all identified requirements. Inspired by prior research on TREC Total Recall Tracks, we introduce the task of Total Recall Question Answering and develop a framework for deep research agents evaluation that satisfies the identified criteria. Our framework constructs single-answer, total recall queries with precise evaluation and relevance judgments derived from a structured knowledge base paired with a text corpus, enabling large-scale data construction. Using this framework, we build TRQA, a deep research benchmark constructed from Wikidata-Wikipedia as a real-world source and a synthetically generated e-commerce knowledge base and corpus to mitigate the effects of data contamination. We benchmark the collection with representative retriever and deep research models and establish baseline retrieval and end-to-end results for future comparative evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。