评估大模型生成医学文献检索式的效果与优化方法
Reassessing Large Language Model Boolean Query Generation for Systematic Reviews
- 复现两篇关键研究并补全验证、格式、提示示例等缺失环节
- 不同模型和提示设计效果差异显著,引导式提示需精选种子文献
- 强调模型与提示的协同优化,适合医学信息检索研究者参考
系统性综述是医学领域最高级别的证据形式,其核心步骤是构建复杂的布尔检索式以获取相关文献。由于手动构建难度高,近期研究探索使用大语言模型(LLMs)辅助生成。早期研究Wang等使用ChatGPT,Staudinger等则在可复现性研究中评估多个LLM,但忽略了原始研究中的关键要素:(i)生成检索式的有效性验证,(ii)输出格式约束,(iii)思维链(引导式)提示中示例的选择。因此其结论与原研究存在显著差异。本文系统复现这两项研究,并补全上述缺失环节。结果表明,检索式效果在不同模型和提示设计间差异明显,引导式生成依赖于精心选择的种子文献。总体而言,提示设计与模型选择是成功生成检索式的关键驱动因素。本研究为理解LLMs在布尔检索式生成中的潜力提供了更清晰的认知,并强调了模型与提示特定优化的重要性。系统性综述本身的复杂性增加了方法开发与复现的挑战,也凸显了该领域可复现性研究的重要价值。
原文摘要 · Abstract (English)
Systematic reviews are comprehensive literature reviews that address highly focused research questions and represent the highest form of evidence in medicine. A critical step in this process is the development of complex Boolean queries to retrieve relevant literature. Given the difficulty of manually constructing these queries, recent efforts have explored Large Language Models (LLMs) to assist in their formulation. One of the first studies,Wang et al., investigated ChatGPT for this task, followed by Staudinger et al., which evaluated multiple LLMs in a reproducibility study. However, the latter overlooked several key aspects of the original work, including (i) validation of generated queries, (ii) output formatting constraints, and (iii) selection of examples for chain-of-thought (Guided) prompting. As a result, its findings diverged significantly from the original study. In this work, we systematically reproduce both studies while addressing these overlooked factors. Our results show that query effectiveness varies significantly across models and prompt designs, with guided query formulation benefiting from well-chosen seed studies. Overall, prompt design and model selection are key drivers of successful query formulation. Our findings provide a clearer understanding of LLMs' potential in Boolean query generation and highlight the importance of model- and prompt-specific optimisations. The complex nature of systematic reviews adds to challenges in both developing and reproducing methods but also highlights the importance of reproducibility studies in this domain.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。