arXiv:2502.15702cs.IRcs.AI2025-02综述被引 4

用大模型自动做系统综述,GPT-4在筛选和数据提取上表现最好。

Large language models streamline automated systematic review: A preliminary study

  • 用三款大模型完成四类综述任务,以真实文献为标准评估
  • GPT-4在文献筛选和数据提取上准确率显著领先
  • 适合研究者快速辅助完成系统综述,降低人工成本

大型语言模型(LLMs)在自然语言处理中展现潜力,有望实现系统综述自动化。本研究评估了GPT-4、Claude-3和Mistral 8x7B在四项系统综述任务中的表现:研究设计制定、检索策略开发、文献筛选和数据提取。基于已有系统综述的参考标准,包括标准PICO(人群、干预、对照、结局)设计、入选标准及20篇参考文献数据,三位评审员使用5分李克特量表评估研究设计与入选标准的质量,涵盖准确性、完整性、相关性、一致性及整体表现。其他任务以是否与参考标准一致判定为准确。检索策略评估准确性与召回效果;筛选准确率分别针对标题摘要和全文;数据提取在1,120个数据点、共3,360个字段上评估。Claude-3在PICO设计上表现最佳;检索策略中GPT-4与Claude-3准确率相当,优于Mistral;标题摘要筛选中GPT-4准确率最高,其次为Mistral和Claude-3;数据提取方面,GPT-4显著优于其他模型。大模型具备自动化系统综述任务潜力,其中GPT-4在检索策略、文献筛选与数据提取上表现突出,是值得进一步开发验证的研究辅助工具。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have shown promise in natural language processing tasks, with the potential to automate systematic reviews. This study evaluates the performance of three state-of-the-art LLMs in conducting systematic review tasks. We assessed GPT-4, Claude-3, and Mistral 8x7B across four systematic review tasks: study design formulation, search strategy development, literature screening, and data extraction. Sourced from a previously published systematic review, we provided reference standard including standard PICO (Population, Intervention, Comparison, Outcome) design, standard eligibility criteria, and data from 20 reference literature. Three investigators evaluated the quality of study design and eligibility criteria using 5-point Liker Scale in terms of accuracy, integrity, relevance, consistency and overall performance. For other tasks, the output is defined as accurate if it is the same as the reference standard. Search strategy performance was evaluated through accuracy and retrieval efficacy. Screening accuracy was assessed for both abstracts screening and full texts screening. Data extraction accuracy was evaluated across 1,120 data points comprising 3,360 individual fields. Claude-3 demonstrated superior overall performance in PICO design. In search strategy formulation, GPT-4 and Claude-3 achieved comparable accuracy, outperforming Mistral. For abstract screening, GPT-4 achieved the highest accuracy, followed by Mistral and Claude-3. In data extraction, GPT-4 significantly outperformed other models. LLMs demonstrate potential for automating systematic review tasks, with GPT-4 showing superior performance in search strategy formulation, literature screening and data extraction. These capabilities make them promising assistive tools for researchers and warrant further development and validation in this field.

大模型系统综述自动化文献筛选

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。