arXiv:2506.12689cs.AIcs.IR2025-06综述被引 14

用多智能体协作生成更严谨的科研综述,提升内容质量和引用准确性。

SciSage: A Multi-Agent Framework for High-Quality Scientific Survey Generation

  • 设计反射式多智能体框架,分层评估综述草稿并协同优化
  • 在文档连贯性上超基准1.73分,引用准确率提升32%
  • 适合需要高效覆盖广度领域、快速检索文献的研究人员

科学文献的快速增长对自动化综述生成工具提出更高要求。现有大模型方法常缺乏深度分析、结构连贯性与可靠引文。为此,我们提出SciSage,一种基于‘写作时反思’范式的多智能体框架。其包含分层反思代理,在大纲、章节和全文层面批判性评估草稿,协同处理查询理解、内容检索与文本优化。同时发布SurveyScope基准,涵盖2020–2025年11个计算机科学领域的46篇高影响力论文,通过严格时效性与引用质量控制筛选。评估显示,SciSage在文档连贯性上优于主流基线(LLM x MapReduce-V2, AutoSurvey)1.73分,引用F1得分提升32%。人工评估结果为3胜7负,但凸显其在主题覆盖面与检索效率上的优势。整体而言,SciSage为研究辅助写作提供了有前景的基础方案。

原文摘要 · Abstract (English)

The rapid growth of scientific literature demands robust tools for automated survey-generation. However, current large language model (LLM)-based methods often lack in-depth analysis, structural coherence, and reliable citations. To address these limitations, we introduce SciSage, a multi-agent framework employing a reflect-when-you-write paradigm. SciSage features a hierarchical Reflector agent that critically evaluates drafts at outline, section, and document levels, collaborating with specialized agents for query interpretation, content retrieval, and refinement. We also release SurveyScope, a rigorously curated benchmark of 46 high-impact papers (2020-2025) across 11 computer science domains, with strict recency and citation-based quality controls. Evaluations demonstrate that SciSage outperforms state-of-the-art baselines (LLM x MapReduce-V2, AutoSurvey), achieving +1.73 points in document coherence and +32% in citation F1 scores. Human evaluations reveal mixed outcomes (3 wins vs. 7 losses against human-written surveys), but highlight SciSage's strengths in topical breadth and retrieval efficiency. Overall, SciSage offers a promising foundation for research-assistive writing tools.

科研写作多智能体综述生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。