arXiv:2412.15247cs.CLcs.IR2024-12被引 38

用大模型自动筛文献,效率提升95%且零漏判。

Streamlining Systematic Reviews: A Novel Application of Large Language Models

  • 用提示工程和RAG技术自动化标题摘要与全文筛选
  • 排除率99.5%,假阴性率0%,仅78篇需人工复核
  • 适合需要高效完成系统综述的研究者

系统综述是循证指南的重要基础,但文献筛选耗时费力。本文提出并评估了一套基于大语言模型(LLMs)的内部系统,用于自动化标题/摘要及全文筛选,填补了该领域的空白。以维生素D与跌倒的系统综述(共14,439篇文献)为例,该系统采用提示工程进行标题/摘要筛选,使用检索增强生成(RAG)进行全文筛选。结果表明,文章排除率(AER)达99.5%,特异性99.6%,假阴性率(FNR)为0%,阴性预测值(NPV)达100%。筛选后仅剩78篇需人工审核,包含传统方法识别出的全部20篇,人工筛查时间减少95.5%。相比之下,商业工具Rayyan在标题/摘要阶段的AER为72.1%,FNR为5%;降低其纳入阈值虽可将FNR降至0%,但增加筛查时间。该系统覆盖双阶段筛选,显著优于Rayyan和传统方法,总筛查时间缩短至25.5小时,同时保持高精度。研究显示,大模型在系统综述流程中具有变革潜力,尤其适用于缺乏自动化工具的全文筛选环节。

原文摘要 · Abstract (English)

Systematic reviews (SRs) are essential for evidence-based guidelines but are often limited by the time-consuming nature of literature screening. We propose and evaluate an in-house system based on Large Language Models (LLMs) for automating both title/abstract and full-text screening, addressing a critical gap in the literature. Using a completed SR on Vitamin D and falls (14,439 articles), the LLM-based system employed prompt engineering for title/abstract screening and Retrieval-Augmented Generation (RAG) for full-text screening. The system achieved an article exclusion rate (AER) of 99.5%, specificity of 99.6%, a false negative rate (FNR) of 0%, and a negative predictive value (NPV) of 100%. After screening, only 78 articles required manual review, including all 20 identified by traditional methods, reducing manual screening time by 95.5%. For comparison, Rayyan, a commercial tool for title/abstract screening, achieved an AER of 72.1% and FNR of 5% when including articles Rayyan considered as undecided or likely to include. Lowering Rayyan's inclusion thresholds improved FNR to 0% but increased screening time. By addressing both screening phases, the LLM-based system significantly outperformed Rayyan and traditional methods, reducing total screening time to 25.5 hours while maintaining high accuracy. These findings highlight the transformative potential of LLMs in SR workflows by offering a scalable, efficient, and accurate solution, particularly for the full-text screening phase, which has lacked automation tools.

系统综述大模型应用文献筛选自动化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。