LUMEN用11个智能体自动完成系统综述,成本仅22.65美元,准确率达100%。
LUMEN: Cost-Transparent Multi-Agent Pipeline for Automated Systematic Review and Meta-Analysis

- 11个专用LLM智能体分阶段协作,按需路由模型提升效率
- 全程成本19-29美元,筛查与提取占总支出超80%
- 首次揭示多智能体架构在不同阶段的性价比差异,适合医学研究者
系统综述与荟萃分析(SR/MA)是循证医学的金标准,但传统方法耗时约67周且依赖专家。现有大语言模型在单个阶段表现优异:筛选(otto-SR:96.7%敏感度)、提取(Gartlehner等:91.0%准确率)、检索(TrialMind:0.83召回率),但尚无研究量化端到端流程的实际成本、成本分布及架构对质量-成本权衡的影响。我们提出LUMEN,一个开源多智能体流水线,使用11个专用LLM智能体自动化六阶段SR/MA流程,并实现智能模型路由。在七项数据集上评估:五项自建领域综述(精神科、心理学、外科、疫苗学、心脏病学)和两项SYNERGY筛选基准。13个可比结果中,LUMEN与既往荟萃分析方向一致率为100%,同质研究设计下效应值误差小于1%。核心贡献在于首次实证揭示此类系统的成本与运行特征:完整综述成本为19至29美元(中位数22.65美元),其中标题摘要筛选与数据提取合计占比超过80%。三组提取消融实验表明阶段依赖型架构反转:多智能体设计虽降低筛选性能,但对提取至关重要,相较单模型方案产出5.7倍可纳入分析,且消除临床危险的方向性错误。两数据集筛选基准显示模型排序具有领域依赖性,不可跨主题迁移。所有代码与成本日志均公开可用。
原文摘要 · Abstract (English)
Systematic reviews and meta-analyses (SR/MA) remain the gold standard for evidence synthesis, yet completing one typically requires 67 weeks and substantial expert effort. Recent large language model (LLM) systems have demonstrated strong performance on individual SR phases - screening (otto-SR: 96.7% sensitivity), extraction (Gartlehner et al.: 91.0% accuracy), and search (TrialMind: 0.83 recall) - but no study has reported what it actually costs to run an end-to-end pipeline, how cost distributes across phases, or how architectural choices affect the cost-quality trade-off. We present LUMEN, an open-source multi-agent pipeline that automates six SR/MA phases using 11 specialized LLM agents with deliberate model routing. We evaluate LUMEN on seven datasets: five self-conducted domain reviews (psychiatry, psychology, surgery, vaccinology, cardiology) and two SYNERGY screening benchmarks. Across 13 ground-truth-comparable outcomes, LUMEN achieves 100% directional agreement with published meta-analyses, with effect sizes within 1% for homogeneous study designs. The primary contribution is the first empirical cost and operational characterization of such a pipeline: a complete review costs 19 to 29 USD (median 22.65 USD), with title-abstract screening and data extraction together dominating expenditure. A three-arm extraction ablation reveals a phase-dependent architecture reversal: multi-agent design hurts screening but is essential for extraction, producing 5.7x more poolable analyses than single-model alternatives while eliminating clinically dangerous direction errors. A two-dataset screening benchmark demonstrates that model ranking is domain-dependent and not transferable across review topics. All code and cost logs are publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。