构建可验证的深度研究任务集,自动演化复杂问题
From Simple QA to Deep Research: A Verifiable Benchmark Constructed through Iterative Task Evolution

- 用迭代式探索-形式化-挑战管道自动生成研究任务
- 含500个任务,覆盖31主题,支持三类查询形式
- 任务以有向无环图结构呈现,支持精准评估
深度研究基准需具备专家级任务与基于领域知识的可靠评估。现有基准多依赖专家撰写或现有人类材料,而完全自动化构建难以保证一致性和可追溯性验证。为弥补这一缺口,我们提出一个包含500个深度研究任务的可验证基准,涵盖31个主题和10大类别,设计了三种查询形式以探测深度研究所需的不同能力。该基准通过迭代的Explorer-Formalizer-Challenger管道自动构建,逐步将简单问题演化为复杂研究任务。每个任务以原子步骤构成的有向无环图(DAG)及对应检查点表示,使查询、DAG与评分标准协同演进。实验表明,该基准能清晰区分模型与查询类型,其基于事实的逐点评分标准支持细粒度、人类对齐且稳定的评估。数据、实现与结果均公开可用。
原文摘要 · Abstract (English)
Deep research benchmarks require expert-level tasks and reliable evaluation grounded in task-specific knowledge. Existing benchmarks rely heavily on expert authoring or pre-existing human-authored materials, while fully automatic construction struggles to ensure consistent and traceable verification. To address this gap, we introduce a verifiable benchmark of 500 deep research tasks spanning 31 topics and 10 major categories, with three query forms designed to probe complementary capabilities required for deep research. The benchmark is constructed automatically using an iterative Explorer-Formalizer-Challenger pipeline that progressively transforms simple questions into deep research tasks. Each task is represented as a directed acyclic graph (DAG) of atomic steps and associated checkpoints, enabling the query, DAG, and rubrics to evolve together in a controlled manner. Experiments demonstrate that the benchmark clearly discriminates among models and query types, while its fact-grounded pointwise rubrics enable fine-grained, human-aligned, and stable evaluation. Our data, implementation, and results are publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。