用大模型零样本排序文献摘要,提升系统综述筛选效率
LGAR: Zero-Shot LLM-Guided Neural Ranking for Abstract Screening in Systematic Literature Reviews
- 用大模型评分+密集重排实现零样本文献相关性排序
- 在57个医学领域综述上比现有方法高5-10个百分点的准确率
- 首次提供完整纳入排除标准数据集,适合医疗领域研究者
科学文献快速增长,使跟踪最新进展变得困难。系统性文献综述(SLRs)旨在识别并评估某一主题的所有相关论文。在检索出候选论文后,摘要筛选阶段确定初步相关性。目前基于大语言模型(LLMs)的摘要筛选方法主要集中在二分类任务;现有的基于问答(QA)的排序方法存在误差传播问题。大语言模型为评估系统综述的纳入与排除标准提供了独特机会,但现有基准未充分覆盖这些标准。我们手动提取了57个系统综述(多为医学领域)的纳入排除标准及研究问题,支持方法间的规范比较。此外,我们提出LGAR,一种零样本的LLM引导抽象排序器,由基于大模型的分级相关性评分器和密集重排器组成。大量实验表明,LGAR在平均精度均值上比现有基于QA的方法高出5-10个百分点。代码与数据已公开。
原文摘要 · Abstract (English)
The scientific literature is growing rapidly, making it hard to keep track of the state-of-the-art. Systematic literature reviews (SLRs) aim to identify and evaluate all relevant papers on a topic. After retrieving a set of candidate papers, the abstract screening phase determines initial relevance. To date, abstract screening methods using large language models (LLMs) focus on binary classification settings; existing question answering (QA) based ranking approaches suffer from error propagation. LLMs offer a unique opportunity to evaluate the SLR's inclusion and exclusion criteria, yet, existing benchmarks do not provide them exhaustively. We manually extract these criteria as well as research questions for 57 SLRs, mostly in the medical domain, enabling principled comparisons between approaches. Moreover, we propose LGAR, a zero-shot LLM Guided Abstract Ranker composed of an LLM based graded relevance scorer and a dense re-ranker. Our extensive experiments show that LGAR outperforms existing QA-based methods by 5-10 pp. in mean average precision. Our code and data is publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。