构建实时研究综述评估基准,推动AI生成文献综述能力发展
DeepScholar-Bench: A Live Benchmark and Automated Evaluation for Generative Research Synthesis
- 从arXiv新论文中提取真实研究问题与人工撰写的范例,构建动态评估任务
- 在知识融合、检索质量与可验证性三维度上,当前最佳系统平均得分仅31%
- 开源参考框架提供基线,适合研究AI辅助科研与自动文献综述的团队使用
人类专家的知识研究与综合能力是进步的核心。新一代AI系统旨在通过实时网络检索生成长篇带引用报告,自动化这一过程。然而,现有评估方法面临挑战:问答基准仅关注短答案,而人工标注数据集存在过时和数据污染问题,无法反映真实研究综述的复杂性与动态性。为此,我们提出DeepScholar-bench,一个面向生成式研究综述的实时基准与自动化评估框架。该基准从高质量arXiv论文中提取查询与人工撰写范例,评估真实综述任务——通过检索、整合并引用先前工作生成相关工作章节。其自动化框架从知识综合、检索质量与可验证性三个关键维度全面衡量性能。为进一步推动研究,我们还发布了DeepScholar-ref,一个基于LOTUS框架的开源参考实现,提供强基线表现。使用DeepScholar-bench,我们系统评估了多个开源系统、搜索代理及OpenAI DeepResearch,发现当前系统几何平均分未超过31%,表明该基准仍有巨大提升空间。这凸显了DeepScholar-bench作为推进生成式研究综述能力基础的重要价值。代码与数据已公开于https://github.com/guestrin-lab/deepscholar-bench。
原文摘要 · Abstract (English)
The ability to research and synthesize knowledge is central to human expertise and progress. A new class of AI systems--designed for generative research synthesis--aims to automate this process by retrieving information from the live web and producing long-form, cited reports. Yet, evaluating such systems remains an open challenge: existing question-answering benchmarks focus on short, factual answers, while expert-curated datasets risk staleness and data contamination. Neither captures the complexity and evolving nature of real research synthesis tasks. We introduce DeepScholar-bench, a live benchmark and automated evaluation framework for generative research synthesis. DeepScholar-bench draws queries and human-written exemplars from recent, high-quality ArXiv papers and evaluates a real synthesis task: generating a related work section by retrieving, synthesizing, and citing prior work. Our automated framework holistically measures performance across three key dimensions--knowledge synthesis, retrieval quality, and verifiability. To further future work, we also contribute DeepScholar-ref, a simple, open-source reference pipeline, which is implemented on the LOTUS framework and provides a strong baseline. Using DeepScholar-bench, we systematically evaluate prior open-source systems, search agents with strong models, OpenAI's DeepResearch, and DeepScholar-ref. We find DeepScholar-bench is far from saturated: no system surpasses a geometric mean of $31\%$ across all metrics. These results highlight both the difficulty and importance of DeepScholar-bench as a foundation for advancing AI systems capable of generative research synthesis. We make our benchmark code and data available at https://github.com/guestrin-lab/deepscholar-bench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。