用LLM提升文献筛选效率,发现双模型协作比单模型更准更可信。
LLM-Assisted Abstract Screening with OLIVER: Evaluating Calibration and Single-Model vs. Actor-Critic Configurations in Literature Reviews
- 采用双模型协作框架(actor-critic)提升筛选判断力
- 单模型配置下准确率高但校准差,假阳性多
- 新方法在两大文献综述中显著降低校准误差,适合大规模筛选
近期研究显示大语言模型(LLMs)可加速文献筛选,但以往评估多基于早期LLMs、标准化Cochrane综述、单一模型设置及以准确率为唯一指标,对泛化性、配置影响和校准能力关注不足。我们开发了OLIVER(用于综述的优化LLM纳入与筛选引擎),一个开源的LLM辅助摘要筛选流程。在两个非Cochrane系统综述中评估多个现代LLMs,分别在全文筛选和最终纳入阶段使用准确率、AUC和校准指标进行评估。此外,测试了结合两个轻量级模型的演员-评论家筛选框架,采用三种聚合规则。结果显示,单个模型表现差异显著:在较小的Review 1(821篇摘要,63篇最终纳入)中,部分模型虽有高敏感度但伴随大量假阳性且校准差;在更大的Review 2(7741篇摘要,71篇最终纳入)中,多数模型特异性强但召回率低,提示提示设计影响召回。所有单模型配置均存在持续性的校准薄弱问题,尽管整体准确率高。而演员-评论家框架在两个综述中均提升判别能力,显著降低校准误差,提高AUC。结论表明,尽管LLMs有望加速筛选,但单模型性能高度依赖于综述特征、提示方式,且校准有限;而演员-评论家框架在保持计算高效的同时,提升分类质量与置信度可靠性,支持低成本大规模筛选。
原文摘要 · Abstract (English)
Introduction: Recent work suggests large language models (LLMs) can accelerate screening, but prior evaluations focus on earlier LLMs, standardized Cochrane reviews, single-model setups, and accuracy as the primary metric, leaving generalizability, configuration effects, and calibration largely unexamined. Methods: We developed OLIVER (Optimized LLM-based Inclusion and Vetting Engine for Reviews), an open-source pipeline for LLM-assisted abstract screening. We evaluated multiple contemporary LLMs across two non-Cochrane systematic reviews and performance was assessed at both the full-text screening and final inclusion stages using accuracy, AUC, and calibration metrics. We further tested an actor-critic screening framework combining two lightweight models under three aggregation rules. Results: Across individual models, performance varied widely. In the smaller Review 1 (821 abstracts, 63 final includes), several models achieved high sensitivity for final includes but at the cost of substantial false positives and poor calibration. In the larger Review 2 (7741 abstracts, 71 final includes), most models were highly specific but struggled to recover true includes, with prompt design influencing recall. Calibration was consistently weak across single-model configurations despite high overall accuracy. Actor-critic screening improved discrimination and markedly reduced calibration error in both reviews, yielding higher AUCs. Discussion: LLMs may eventually accelerate abstract screening, but single-model performance is highly sensitive to review characteristics, prompting, and calibration is limited. An actor-critic framework improves classification quality and confidence reliability while remaining computationally efficient, enabling large-scale screening at low cost.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。