用多智能体协作生成更可靠易读的文献综述。
LiRA: A Multi-Agent Framework for Reliable and Readable Literature Review Generation
- 分角色智能体协同完成文献综述的提纲、写作、编辑与评审。
- 在两个数据集上优于AutoSurvey等基线,接近人工写作质量。
- 无需领域微调,适合科研人员快速撰写高质量综述。
科学论文的快速增长使文献综述保持全面和更新愈发困难。尽管已有研究聚焦于自动化检索与筛选,系统性综述的撰写阶段仍缺乏深入探索,尤其在可读性和事实准确性方面。为此,我们提出LiRA(Literature Review Agents),一种模拟人类文献综述流程的多智能体协作框架。LiRA采用专门负责内容提纲、子章节撰写、编辑和审阅的智能体,生成连贯且全面的综述文章。在SciReviewGen和自有ScienceDirect数据集上的评估显示,LiRA在写作质量和引用准确性上超越AutoSurvey和MASS-Survey等基线模型,同时与人工撰写综述保持相近相似度。我们进一步在真实文档检索场景中测试其表现,并评估了不同评审模型下的鲁棒性。结果表明,即使无领域特定微调,基于智能体的LLM工作流也能显著提升自动化科学写作的可靠性与可用性。
原文摘要 · Abstract (English)
The rapid growth of scientific publications has made it increasingly difficult to keep literature reviews comprehensive and up-to-date. Though prior work has focused on automating retrieval and screening, the writing phase of systematic reviews remains largely under-explored, especially with regard to readability and factual accuracy. To address this, we present LiRA (Literature Review Agents), a multi-agent collaborative workflow which emulates the human literature review process. LiRA utilizes specialized agents for content outlining, subsection writing, editing, and reviewing, producing cohesive and comprehensive review articles. Evaluated on SciReviewGen and a proprietary ScienceDirect dataset, LiRA outperforms current baselines such as AutoSurvey and MASS-Survey in writing and citation quality, while maintaining competitive similarity to human-written reviews. We further evaluate LiRA in real-world scenarios using document retrieval and assess its robustness to reviewer model variation. Our findings highlight the potential of agentic LLM workflows, even without domain-specific tuning, to improve the reliability and usability of automated scientific writing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。