arXiv:2411.02451cs.CLcs.DL2024-11被引 17

用大模型组批量筛选论文摘要,比人还准还快。

High-performance automated abstract screening with large language model ensembles

  • 用多个大模型组合零样本分类,自动判断论文是否纳入综述。
  • 在800篇样本中,模型灵敏度达1.000,高于人类的0.775。
  • 适合需要高效处理海量文献的科研人员和系统综述团队。

大语言模型(LLMs)在文本处理与理解任务中表现优异。系统综述中的摘要筛选是耗时繁重的工作,需对大量研究逐一应用纳入与排除标准。本研究在《Cochrane图书馆》整期文献中测试了GPT-3.5 Turbo、GPT-4 Turbo、GPT-4o、Llama 3 70B、Gemini 1.5 Pro和Claude Sonnet 3.5等6种模型,在零样本二分类任务中评估其准确性。在800条记录的子集上,优化提示策略后,模型在灵敏度(最高1.000)、精确率(最高0.927)和平衡准确率(最高0.904)上均优于人类(最高分别为0.775、0.911、0.865)。最优模型-提示组合应用于全部119,691条检索结果,灵敏度保持在0.756–1.000之间,但精确率下降至0.004–0.096。66种模型-人类及模型-模型集成方案实现了完美灵敏度,最高精确率为0.458,且大样本下性能下降更小。不同综述间表现差异显著,强调部署前需进行领域特定验证。大模型可显著降低系统综述的人力成本,同时维持或提升准确性和灵敏度。系统综述是证据综合的基础,涵盖循证医学等多个学科,大模型有望提升研究效率与质量。

原文摘要 · Abstract (English)

Large language models (LLMs) excel in tasks requiring processing and interpretation of input text. Abstract screening is a labour-intensive component of systematic review involving repetitive application of inclusion and exclusion criteria on a large volume of studies identified by a literature search. Here, LLMs (GPT-3.5 Turbo, GPT-4 Turbo, GPT-4o, Llama 3 70B, Gemini 1.5 Pro, and Claude Sonnet 3.5) were trialled on systematic reviews in a full issue of the Cochrane Library to evaluate their accuracy in zero-shot binary classification for abstract screening. Trials over a subset of 800 records identified optimal prompting strategies and demonstrated superior performance of LLMs to human researchers in terms of sensitivity (LLM-max = 1.000, human-max = 0.775), precision (LLM-max = 0.927, human-max = 0.911), and balanced accuracy (LLM-max = 0.904, human-max = 0.865). The best performing LLM-prompt combinations were trialled across every replicated search result (n = 119,691), and exhibited consistent sensitivity (range 0.756-1.000) but diminished precision (range 0.004-0.096). 66 LLM-human and LLM-LLM ensembles exhibited perfect sensitivity with a maximal precision of 0.458, with less observed performance drop in larger trials. Significant variation in performance was observed between reviews, highlighting the importance of domain-specific validation before deployment. LLMs may reduce the human labour cost of systematic review with maintained or improved accuracy and sensitivity. Systematic review is the foundation of evidence synthesis across academic disciplines, including evidence-based medicine, and LLMs may increase the efficiency and quality of this mode of research.

大模型文献筛选系统综述自动化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。