arXiv:2411.14914cs.IR2024-11被引 28

用大模型生成文献检索关键词,测试其可复现性和通用性。

A Reproducibility and Generalizability Study of Large Language Models for Query Generation

  • 构建自动化流程:用LLM生成布尔查询,从PubMed检索文献并评估结果。
  • ChatGPT生成的查询可复现,但开源模型如Mistral和Zephyr表现差异明显。
  • 揭示了大模型在文献检索中的局限,适合研究自动化综述的学者参考。

系统性文献综述(SLRs)是学术研究的基础,但因文献筛选过程繁琐而耗时。生成式AI与大语言模型(LLMs)有望通过辅助生成有效布尔查询来革新这一流程。本文对使用LLMs生成布尔查询用于系统性综述进行了全面研究,复现并扩展了Wang等与Alaniz等的工作。研究考察了ChatGPT生成查询的可复现性,并对比其与开源模型(如Mistral、Zephyr)的性能,提供更全面的分析。我们构建了一个自动化流水线:基于预定义的LLM生成指定主题的布尔查询,在PubMed中检索相关文献并评估结果。首先验证了使用ChatGPT生成查询的结果是否可复现且一致;随后通过分析开源模型推广结论,评估其生成布尔查询的有效性;最后进行故障分析,识别并讨论使用LLMs生成布尔查询的局限与不足。该研究有助于理解大模型在信息检索任务中的差距与改进方向。研究结果突显了大模型在信息检索与文献综述自动化中的潜力、局限与前景。

原文摘要 · Abstract (English)

Systematic literature reviews (SLRs) are a cornerstone of academic research, yet they are often labour-intensive and time-consuming due to the detailed literature curation process. The advent of generative AI and large language models (LLMs) promises to revolutionize this process by assisting researchers in several tedious tasks, one of them being the generation of effective Boolean queries that will select the publications to consider including in a review. This paper presents an extensive study of Boolean query generation using LLMs for systematic reviews, reproducing and extending the work of Wang et al. and Alaniz et al. Our study investigates the replicability and reliability of results achieved using ChatGPT and compares its performance with open-source alternatives like Mistral and Zephyr to provide a more comprehensive analysis of LLMs for query generation. Therefore, we implemented a pipeline, which automatically creates a Boolean query for a given review topic by using a previously defined LLM, retrieves all documents for this query from the PubMed database and then evaluates the results. With this pipeline we first assess whether the results obtained using ChatGPT for query generation are reproducible and consistent. We then generalize our results by analyzing and evaluating open-source models and evaluating their efficacy in generating Boolean queries. Finally, we conduct a failure analysis to identify and discuss the limitations and shortcomings of using LLMs for Boolean query generation. This examination helps to understand the gaps and potential areas for improvement in the application of LLMs to information retrieval tasks. Our findings highlight the strengths, limitations, and potential of LLMs in the domain of information retrieval and literature review automation.

大模型文献检索自动化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。