用AI生成高质量论文综述,提升结构与引用准确性
SurveyForge: On the Outline Heuristics, Memory-Driven Generation, and Multi-dimensional Evaluation for Automated Survey Writing
- 分析人类综述结构并检索相关论文,生成逻辑清晰的提纲
- 通过学者导航代理获取高质量文献,自动撰写并优化内容
- 构建多维度评估基准,适合需要高效撰写综述的研究者
综述论文在科学研究中至关重要,尤其面对研究文献的快速增长。近期研究开始利用大语言模型(LLM)自动化生成综述以提高效率。然而,当前LLM生成的综述在提纲质量和引用准确性方面仍与人工撰写的有显著差距。为此,我们提出SurveyForge,首先通过分析人类撰写综述的逻辑结构,并结合检索到的领域相关论文,生成高质量提纲;随后,借助由学者导航代理从记忆中检索出的高质量论文,自动撰写并迭代优化文章内容。此外,为实现全面评估,我们构建了SurveyBench基准,包含100篇人工撰写的综述论文,用于胜率对比,并从参考文献、提纲和内容质量三个维度评估AI生成的综述。实验表明,SurveyForge优于先前方法如AutoSurvey。
原文摘要 · Abstract (English)
Survey paper plays a crucial role in scientific research, especially given the rapid growth of research publications. Recently, researchers have begun using LLMs to automate survey generation for better efficiency. However, the quality gap between LLM-generated surveys and those written by human remains significant, particularly in terms of outline quality and citation accuracy. To close these gaps, we introduce SurveyForge, which first generates the outline by analyzing the logical structure of human-written outlines and referring to the retrieved domain-related articles. Subsequently, leveraging high-quality papers retrieved from memory by our scholar navigation agent, SurveyForge can automatically generate and refine the content of the generated article. Moreover, to achieve a comprehensive evaluation, we construct SurveyBench, which includes 100 human-written survey papers for win-rate comparison and assesses AI-generated survey papers across three dimensions: reference, outline, and content quality. Experiments demonstrate that SurveyForge can outperform previous works such as AutoSurvey.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。