用多智能体系统实现端到端自动化元分析,减少幻觉错误。
Manalyzer: End-to-end Automated Meta-analysis with Multi-agent System
- 设计多智能体系统,通过工具调用完成文献检索到数据提取全流程。
- 在729篇论文的跨模态基准上,数据点准确率显著优于传统LLM基线。
- 适合需要高效整合大量研究结论的研究人员与领域专家使用。
元分析是一种系统性研究方法,通过整合多个已有研究的数据以得出综合性结论,不仅可克服单个研究的局限性,还能通过数据融合发现新规律。传统元分析流程复杂,包含文献检索、论文筛选和数据提取等多个阶段,需耗费大量人力与时间。尽管基于大模型的方法可加速部分环节,但仍面临论文筛选和数据提取中的幻觉问题。本文提出多智能体系统Manalyzer,通过工具调用实现端到端自动化元分析。其采用混合评审、分层提取、自证机制与反馈校验策略,显著缓解上述两类幻觉。为全面评估性能,我们构建了一个涵盖3个领域的基准,包含729篇论文,覆盖文本、图像与表格模态,含超过10,000个数据点。大量实验表明,Manalyzer在多项元分析任务中显著优于主流大模型基线。
原文摘要 · Abstract (English)
Meta-analysis is a systematic research methodology that synthesizes data from multiple existing studies to derive comprehensive conclusions. This approach not only mitigates limitations inherent in individual studies but also facilitates novel discoveries through integrated data analysis. Traditional meta-analysis involves a complex multi-stage pipeline including literature retrieval, paper screening, and data extraction, which demands substantial human effort and time. However, while LLM-based methods can accelerate certain stages, they still face significant challenges, such as hallucinations in paper screening and data extraction. In this paper, we propose a multi-agent system, Manalyzer, which achieves end-to-end automated meta-analysis through tool calls. The hybrid review, hierarchical extraction, self-proving, and feedback checking strategies implemented in Manalyzer significantly alleviate these two hallucinations. To comprehensively evaluate the performance of meta-analysis, we construct a new benchmark comprising 729 papers across 3 domains, encompassing text, image, and table modalities, with over 10,000 data points. Extensive experiments demonstrate that Manalyzer achieves significant performance improvements over the LLM baseline in multi meta-analysis tasks. Project page: https://black-yt.github.io/meta-analysis-page/ .
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。