用可复现的提示编译替代手工撰写,提升AI辅助文献综述可靠性
Compiling Prompts, Not Crafting Them: A Reproducible Workflow for AI-Assisted Evidence Synthesis
- 将声明式提示优化技术引入文献综述,实现自动化提示调优
- 构建包含测试集与任务定义的结构化工作流,确保结果可复现
- 提供开源代码模板,适合需要高透明度的研究者使用
大型语言模型(LLMs)在加速系统性文献综述(SLRs)方面潜力巨大,但当前方法多依赖脆弱的手工编写提示,损害了可靠性和可复现性,削弱了科学界对AI辅助证据综合的信心。本文借鉴近期为通用LLM应用开发的声明式提示优化技术,证明其在SLR自动化领域的适用性。研究提出一个领域特定的结构化框架,将任务声明、测试套件和自动提示调优嵌入可复现的SLR流程中。这些新兴方法被转化为具体可执行的蓝图,并附有可运行代码示例,使研究人员能够构建符合透明度与严谨性原则的可验证LLM管道。这是此类方法首次应用于SLR流程。
原文摘要 · Abstract (English)
Large language models (LLMs) offer significant potential to accelerate systematic literature reviews (SLRs), yet current approaches often rely on brittle, manually crafted prompts that compromise reliability and reproducibility. This fragility undermines scientific confidence in LLM-assisted evidence synthesis. In response, this work adapts recent advances in declarative prompt optimisation, developed for general-purpose LLM applications, and demonstrates their applicability to the domain of SLR automation. This research proposes a structured, domain-specific framework that embeds task declarations, test suites, and automated prompt tuning into a reproducible SLR workflow. These emerging methods are translated into a concrete blueprint with working code examples, enabling researchers to construct verifiable LLM pipelines that align with established principles of transparency and rigour in evidence synthesis. This is a novel application of such approaches to SLR pipelines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。