用大模型自动完成文献综述,提升效率并验证可行性。
The emergence of Large Language Models (LLM) as a tool in literature reviews: an LLM automated systematic review
- 用GPT-4o等大模型辅助文献筛选与数据提取
- 3788篇文献中筛选出172篇合格研究,自动化率达73.2%
- 适合科研人员快速构建系统性综述
目的:本研究旨在总结大语言模型(LLM)在科学综述撰写过程中的应用现状。我们分析了可实现自动化的综述阶段,并评估该领域的前沿研究项目。方法:2024年6月,人工审阅者在PubMed、Scopus、Dimensions和Google Scholar数据库中进行检索。文献筛选与数据提取通过Covidence平台结合使用OpenAI gpt-4o模型的LLM插件完成。使用ChatGPT清洗提取数据,并生成本文图表代码;ChatGPT与Scite.ai协助撰写除方法与讨论外的所有部分。结果:共检索到3,788篇文献,最终纳入172项研究。基于ChatGPT和GPT的模型在综述自动化中占据主导地位(n=126,占73.2%)。虽有大量自动化项目,但仅有26项(15.1%)实际采用大模型完成综述。多数研究聚焦于特定阶段自动化,如文献搜索(n=60,34.9%)与数据提取(n=54,31.4%)。对比分析显示,GPT类模型在数据提取上表现更优,平均精确率83.0%(SD=10.4),召回率86.0%(SD=9.8);但在标题与摘要筛选阶段略逊,准确率77.3%(SD=13.0)。结论:本研究揭示了大量基于大模型的综述自动化项目,结果令人鼓舞,预计未来大模型将深刻改变科学综述的撰写方式。
原文摘要 · Abstract (English)
Objective: This study aims to summarize the usage of Large Language Models (LLMs) in the process of creating a scientific review. We look at the range of stages in a review that can be automated and assess the current state-of-the-art research projects in the field. Materials and Methods: The search was conducted in June 2024 in PubMed, Scopus, Dimensions, and Google Scholar databases by human reviewers. Screening and extraction process took place in Covidence with the help of LLM add-on which uses OpenAI gpt-4o model. ChatGPT was used to clean extracted data and generate code for figures in this manuscript, ChatGPT and Scite.ai were used in drafting all components of the manuscript, except the methods and discussion sections. Results: 3,788 articles were retrieved, and 172 studies were deemed eligible for the final review. ChatGPT and GPT-based LLM emerged as the most dominant architecture for review automation (n=126, 73.2%). A significant number of review automation projects were found, but only a limited number of papers (n=26, 15.1%) were actual reviews that used LLM during their creation. Most citations focused on automation of a particular stage of review, such as Searching for publications (n=60, 34.9%), and Data extraction (n=54, 31.4%). When comparing pooled performance of GPT-based and BERT-based models, the former were better in data extraction with mean precision 83.0% (SD=10.4), and recall 86.0% (SD=9.8), while being slightly less accurate in title and abstract screening stage (Maccuracy=77.3%, SD=13.0). Discussion/Conclusion: Our LLM-assisted systematic review revealed a significant number of research projects related to review automation using LLMs. The results looked promising, and we anticipate that LLMs will change in the near future the way the scientific reviews are conducted.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。