用智能体+评分标准自动写学术综述,质量媲美人类专家。
ARISE: Agentic Rubric-Guided Iterative Survey Engine for Automated Scholarly Paper Generation
- 分角色智能体协作生成论文,按评分标准迭代优化。
- 综合评分达92.48,内容完整度、准确性和格式均优于现有方法。
- 适合需要快速产出高质量综述的研究者和团队使用。
学术文献的快速增长给撰写全面高质量的综述带来巨大挑战。近年来,智能体系统在自动化文献综述、整合与迭代优化等传统上需人工完成的任务中展现出巨大潜力。然而,现有自动化综述生成方案普遍存在质量控制不足、格式不规范、难以适应迭代反馈等问题,而这些正是学术写作的核心要素。为此,我们提出ARISE——一种基于评分标准引导的智能体式迭代综述生成引擎,用于自动化生成并持续优化学术综述论文。ARISE采用模块化架构,包含多个专用大语言模型智能体,分别模拟主题拓展、引用整理、文献摘要、稿件撰写及基于同行评审的评估等学术角色。其核心是基于结构化行为锚定评分表的迭代优化循环,多个评审智能体独立评估稿件并提供合成反馈,持续提升内容质量。在与先进自动化系统及近期人工撰写的综述对比的实验中,ARISE表现出色,平均评分达到92.48,在全面性、准确性、格式规范性和整体学术严谨性等指标上均显著优于基线方法。所有代码、评价标准及生成结果已开源,详见 https://github.com/ziwang11112/ARISE。
原文摘要 · Abstract (English)
The rapid expansion of scholarly literature presents significant challenges in synthesizing comprehensive, high-quality academic surveys. Recent advancements in agentic systems offer considerable promise for automating tasks that traditionally require human expertise, including literature review, synthesis, and iterative refinement. However, existing automated survey-generation solutions often suffer from inadequate quality control, poor formatting, and limited adaptability to iterative feedback, which are core elements intrinsic to scholarly writing. To address these limitations, we introduce ARISE, an Agentic Rubric-guided Iterative Survey Engine designed for automated generation and continuous refinement of academic survey papers. ARISE employs a modular architecture composed of specialized large language model agents, each mirroring distinct scholarly roles such as topic expansion, citation curation, literature summarization, manuscript drafting, and peer-review-based evaluation. Central to ARISE is a rubric-guided iterative refinement loop in which multiple reviewer agents independently assess manuscript drafts using a structured, behaviorally anchored rubric, systematically enhancing the content through synthesized feedback. Evaluating ARISE against state-of-the-art automated systems and recent human-written surveys, our experimental results demonstrate superior performance, achieving an average rubric-aligned quality score of 92.48. ARISE consistently surpasses baseline methods across metrics of comprehensiveness, accuracy, formatting, and overall scholarly rigor. All code, evaluation rubrics, and generated outputs are provided openly at https://github.com/ziwang11112/ARISE
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。