arXiv:2510.01283cs.CL2025-10综述

为深度研究工具设计评估标准,用学术综述写作验证其效果

Evaluation Sheet for Deep Research: A Use Case for Academic Survey Writing

  • 提出针对深度研究工具的评估表,聚焦信息提取与报告生成能力
  • 对比OpenAI和Google深搜生成的学术综述,发现存在显著质量差距
  • 适合关注AI辅助科研、论文写作工具评估的研究者参考

具备代理能力的大语言模型可实现无需人工干预的知识密集型任务。以深度研究为例,该工具具备网络浏览、信息提取和生成多页报告的能力。本文提出一种用于评估深度研究工具性能的评估表,并以学术综述写作作为具体应用案例,对输出报告进行评估。实验结果表明,建立精细的评估标准至关重要。在对OpenAI Deep Search和Google Deep Search生成学术综述的评估中,揭示了搜索引擎与独立深度研究工具之间存在的巨大差距,尤其体现在对目标领域的全面性和准确性表现不足。

原文摘要 · Abstract (English)

Large Language Models (LLMs) powered with argentic capabilities are able to do knowledge-intensive tasks without human involvement. A prime example of this tool is Deep research with the capability to browse the web, extract information and generate multi-page reports. In this work, we introduce an evaluation sheet that can be used for assessing the capability of Deep Research tools. In addition, we selected academic survey writing as a use case task and evaluated output reports based on the evaluation sheet we introduced. Our findings show the need to have carefully crafted evaluation standards. The evaluation done on OpenAI`s Deep Search and Google's Deep Search in generating an academic survey showed the huge gap between search engines and standalone Deep Research tools, the shortcoming in representing the targeted area.

深度研究LLM评估学术写作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。