arXiv:2608.18034cs.CV2026-08综述

用智能体闭环系统自动生成高质量学术综述,提升写作与验证效率。

Deep Academic Survey: Stateful Agentic Closed-Loop Paradigm for Academic Survey Automation

论文配图:Deep Academic Survey: Stateful Agentic Closed-Loop Paradigm for Academic Survey Automation
图 1 · 摘自论文原文
  • 构建可复用的论文分析与主题化写作分离的智能体框架
  • 在30个主题上综合得分4.34,优于最强对手4.03
  • 适合需要高效、可靠综述生成的研究者与期刊编辑

学术综述在整理快速扩张的文献中起核心作用,但其构建需大量论文分析、知识组织、细粒度引证支持及可靠文稿整合。现有深度研究与自动化综述系统仅覆盖部分流程,且缺乏对论文理解、文献组织、证据支撑写作与文稿验证的统一状态管理。本文提出DAS,一种面向出版的有状态智能体闭环框架。其核心思想是将可复用的论文分析与特定主题的文稿构建相分离。DAS基于DAS-2M——一个包含约两百万篇论文的动态更新元数据湖,其智能体通过候选引导的分类规划、反向论文到章节映射及层级化论点与引用规划,维护显式的文献、组织、写作与定稿状态。语义审查仅激活受影响的写作状态进行修复与重评,形成确定性验证的局部闭环。我们还引入DAS-Bench(30个主题)与DAS-Eval,通过16项指标评估引证质量、分类合成、层级论述与文稿可靠性。在全部30个主题中,DAS四项维度均排名第一,平均分4.34,领先最强对手4.03;在21个共享计算机科学主题子集上保持相同排名。盲评专家在30个主题中27次偏好DAS胜过Naive RAG,19次胜过AutoSurvey。

原文摘要 · Abstract (English)

Academic surveys play a central role in organizing rapidly expanding scholarly literature, yet their construction requires extensive paper analysis, coherent knowledge organization, fine-grained citation support, and reliable manuscript assembly. Existing Deep Research and automated survey generation systems address parts of this process, but typically do not coordinate paper understanding, literature organization, evidence-grounded drafting, and manuscript validation through a shared, revisable state. We introduce DAS, a stateful agentic framework for generating publication-oriented academic surveys. Its key idea is to separate reusable paper analysis from topic-specific manuscript construction. DAS builds on DAS-2M, a dynamically updated metadata lake containing survey-oriented representations of approximately two million papers. Its agents maintain explicit literature, organization, writing, and finalization states through candidate-grounded taxonomy planning, reverse paper-to-section routing, and hierarchical claim and citation planning. Semantic review reactivates only the affected writing states for repair and reevaluation, forming a scoped closed loop with deterministic validation. We further introduce DAS-Bench, a 30-topic benchmark, together with DAS-Eval, which assesses scholarly citation quality, taxonomic synthesis, hierarchical discourse, and manuscript assembly reliability through 16 criteria. Among systems evaluated on all 30 topics, DAS achieves the highest average in all four dimensions, with an overall score of 4.34 compared with 4.03 for the strongest competitor, and the same ordering is preserved on the matched 21-topic CS subset. Blinded expert evaluation further prefers DAS to Naive RAG on 27 of 30 topics and to AutoSurvey on 19 of 21 shared CS topics. The project page is available at https://zhikaixu24.github.io/projects/DAS/.

学术综述智能体系统闭环生成文献管理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。