arXiv:2607.24991cs.SEcs.AI2026-07

为用生成式AI做文献综述提供可操作的指导框架

Preliminary Guidelines for Using and Evaluating GenAI Tools to Support Systematic Literature Reviews

  • 提出GUEST框架,规范使用AI开展文献综述的流程
  • 强调人类监督必要性,指出当前AI无法独立完成系统综述
  • 适合需提升效率又保证质量的软件工程研究者

背景:生成式AI和大语言模型在软件工程等领域的学术任务中日益普及,包括系统文献综述(SLRs)。尽管能摘要文本,但尚无保障其满足系统综述所需的严谨性、可靠性和透明度。目标:为计划使用生成式AI进行系统文献综述的研究者,或开展评估生成式AI支持综述任务实证研究的研究者提供支持。方法:首先开展快速文献回顾,识别已有关于评估和使用生成式AI及大语言模型支持系统文献综述的指南;其次结合思想实验、文献相关指引及作者自身执行系统文献综述与工具评估的经验,提出生成式AI在系统文献综述中的使用与评估建议。结果:讨论了研究人员在评估生成式AI用于系统文献综述时面临的问题,识别并解释了在规划、执行和报告使用生成式AI的系统文献综述及生成式AI工具评估时应考虑的过程问题。最终总结为一套过程性建议,命名为GUEST(GenAI Use and Evaluation in SLR Tasks)。结论:我们主张生成式AI需要人类监督,目前尚不能实现无监督的系统研究。然而,它为部分重复性任务提供了成本效益高的辅助,并可用于复杂任务的额外验证。我们的GUEST建议有助于软件工程研究者在使用生成式AI时开展并报告可信的系统文献综述,同时推动严谨的独立评估研究。

原文摘要 · Abstract (English)

Context: Generative AI (GenAI) and Large Language Models (LLMs) are increasingly used for academic tasks in software engineering and beyond, including systematic literature reviews (SLRs). However, while capable of summarizing text, there is no guarantee they can meet the rigour, reliability, and transparency that SLRs require. Objectives: To support researchers intending to conduct SLRs using GenAI or those conducting empirical studies evaluating how well GenAI supports SLR tasks. Methods: First, we conducted a rapid review to identify studies that propose guidelines for evaluating and using GenAI and LLMs to support SLRs. Second, we drew on thought experiments, relevant guidance from the literature, and our own experience conducting SLRs and evaluating tools to develop recommendations for how to use and assess GenAI in the context of SLRs. Results: We discuss the problems researchers face when evaluating GenAI for SLRs. We identify and explain process issues to consider when planning, conducting, and reporting both SLRs using GenAI and evaluations of GenAI tools. Finally, we summarize our results as a set of process recommendations, which we name GUEST (GenAI Use and Evaluation in SLR Tasks). Conclusion: We argue that GenAI requires human oversight and is not currently capable of unsupervised systematic studies. However, it offers the prospect of cost-effective assistance for some repetitive tasks and for additional validation of some complex tasks. Our GUEST recommendations should help software engineering researchers both to conduct and report trustworthy SLRs using GenAI and to provide rigorous independent evaluation studies.

文献综述生成式AI研究方法GUEST框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。