用多智能体系统评估文献综述质量,准确率达84%。
Can Agents Judge Systematic Reviews Like Humans? Evaluating SLRs with LLM-based Multi-Agent System
- 构建基于大模型的多智能体系统,按PRISMA指南自动评估
- 在5篇跨领域综述上与专家评分达成84%一致
- 适合需要高效、标准化文献综述的研究者
系统性文献综述(SLRs)是循证研究的基础,但耗时且易因学科差异导致不一致。本文提出一种基于大语言模型的多智能体系统(MAS)架构,作为文献综述质量评估辅助工具,可自动完成方案验证、方法学评估和主题相关性检查。相比传统单智能体方法,该系统采用符合PRISMA指南的专用智能体协作机制,提升评估的结构化与可解释性。我们在五个来自不同领域的已发表文献综述上进行了初步研究,将系统输出与专家标注的PRISMA评分对比,观察到84%的一致性。尽管结果尚处早期阶段,本工作标志着迈向可扩展、高精度的NLP驱动跨学科综述流程的第一步,展现了其在无领域依赖的知识整合能力,有望简化综述流程。
原文摘要 · Abstract (English)
Systematic Literature Reviews (SLRs) are foundational to evidence-based research but remain labor-intensive and prone to inconsistency across disciplines. We present an LLM-based SLR evaluation copilot built on a Multi-Agent System (MAS) architecture to assist researchers in assessing the overall quality of the systematic literature reviews. The system automates protocol validation, methodological assessment, and topic relevance checks using a scholarly database. Unlike conventional single-agent methods, our design integrates a specialized agentic approach aligned with PRISMA guidelines to support more structured and interpretable evaluations. We conducted an initial study on five published SLRs from diverse domains, comparing system outputs to expert-annotated PRISMA scores, and observed 84% agreement. While early results are promising, this work represents a first step toward scalable and accurate NLP-driven systems for interdisciplinary workflows and reveals their capacity for rigorous, domain-agnostic knowledge aggregation to streamline the review process.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。