用大模型自动构建系统性综述检索策略,提升召回率与可复现性。
Chained Prompting for Better Systematic Review Search Strategies
- 通过链式提示分解综述目标,自动生成结构化PICO元素
- 在LEADSInstruct数据集上实现0.9平均召回率,优于现有方法
- 适合需要高效、高召回检索的科研人员和循证医学实践者
系统性综述依赖精心设计的检索策略以实现全面查全并减少偏倚。传统人工方法虽系统但耗时且易受主观影响,而启发式与自动化技术在缺乏专家干预时召回率普遍偏低。本文提出一种基于大语言模型(LLM)的链式提示工程框架,用于自动构建系统性综述的检索策略。该框架模仿人工设计流程,利用LLM分解综述目标、提取并形式化PICO要素、生成概念表征、扩展术语并合成布尔查询。除查询生成外,其在生成结构化PICO要素方面表现优于现有方法,为高查全率检索奠定基础。在LEADSInstruct数据集子集上的评估显示,该框架达到0.9平均召回率,显著超越现有方法。错误分析进一步凸显精确目标定义与术语对齐在优化检索效果中的关键作用。结果证实,基于LLM的流水线具备生成透明、可复现且高性能检索策略的能力,具有作为可扩展工具支持证据综合与循证实践的潜力。
原文摘要 · Abstract (English)
Systematic reviews require the use of rigorously designed search strategies to ensure both comprehensive retrieval and minimization of bias. Conventional manual approaches, although methodologically systematic, are resource-intensive and susceptible to subjectivity, whereas heuristic and automated techniques frequently under-perform in recall unless supplemented by extensive expert input. We introduce a Large Language Model (LLM)-based chained prompt engineering framework for the automated development of search strategies in systematic reviews. The framework replicates the procedural structure of manual search design while leveraging LLMs to decompose review objectives, extract and formalize PICO elements, generate conceptual representations, expand terminologies, and synthesize Boolean queries. In addition to query construction, the framework exhibits superior performance in generating well-structured PICO elements relative to existing methods, thereby strengthening the foundation for high-recall search strategies. Evaluation on a subset of the LEADSInstruct dataset demonstrates that the framework attains a 0.9 average recall. These results significantly exceed the performance of existing approaches. Error analysis further highlights the critical role of precise objective specification and terminological alignment in optimizing retrieval effectiveness. These findings confirm the capacity of LLM-based pipelines to yield transparent, reproducible, and high-performing search strategies, and highlight their potential as scalable instruments for supporting evidence synthesis and evidence-based practice.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。