arXiv:2510.16091cs.CLcs.AI2025-10综述被引 5

对比不同提示策略与大模型在文献筛选中的表现,找到高效低成本的自动化方案。

Evaluating Prompting Strategies and Large Language Models in Systematic Literature Review Screening: Relevance and Task-Stage Classification

  • 用五种提示方式测试六种大模型,评估其在相关性与任务阶段分类的表现。
  • 思维链+少样本提示效果最稳,零样本召回率最高,但自省提示易过泛且不稳定。
  • 推荐分阶段使用:低成本模型初筛,可疑条目再交由高性能模型复核。

本研究量化了提示策略与大语言模型(LLMs)在系统性文献综述(SLRs)筛选阶段的交互效应。我们评估了六种模型(GPT-4o、GPT-4o-mini、DeepSeek-Chat-V3、Gemini-2.5-Flash、Claude-3.5-Haiku、Llama-4-Maverick)在五类提示(零样本、少样本、思维链CoT、CoT-少样本、自省)下的表现,涵盖相关性分类和六个二级任务,使用准确率、精确率、召回率与F1值进行评估。结果表明存在显著的模型-提示交互效应:CoT-少样本提示在精确率-召回率平衡上最优;零样本在高敏感度筛查中实现最高召回率;自省提示因过度包容和跨模型不稳定性而表现较差。GPT-4o与DeepSeek整体表现稳健,而GPT-4o-mini在成本大幅降低的前提下仍具竞争力。针对相关性分类的成本-性能分析显示,每千篇摘要的绝对差异显著;GPT-4o-mini在各类提示下均保持低耗,且在该模型上使用结构化提示(CoT/CoT-少样本)可实现高F1值,仅需小幅增量成本。建议采用分阶段流程:(1)以低成本模型配合结构化提示进行首轮筛选;(2)仅将边界案例升级至高能力模型处理。研究揭示了大模型在文献筛选中虽不均衡但潜力可观,通过系统分析提示-模型互动,提供了可比基准与任务适配部署的实用指导。

原文摘要 · Abstract (English)

This study quantifies how prompting strategies interact with large language models (LLMs) to automate the screening stage of systematic literature reviews (SLRs). We evaluate six LLMs (GPT-4o, GPT-4o-mini, DeepSeek-Chat-V3, Gemini-2.5-Flash, Claude-3.5-Haiku, Llama-4-Maverick) under five prompt types (zero-shot, few-shot, chain-of-thought (CoT), CoT-few-shot, self-reflection) across relevance classification and six Level-2 tasks, using accuracy, precision, recall, and F1. Results show pronounced model-prompt interaction effects: CoT-few-shot yields the most reliable precision-recall balance; zero-shot maximizes recall for high-sensitivity passes; and self-reflection underperforms due to over-inclusivity and instability across models. GPT-4o and DeepSeek provide robust overall performance, while GPT-4o-mini performs competitively at a substantially lower dollar cost. A cost-performance analysis for relevance classification (per 1,000 abstracts) reveals large absolute differences among model-prompt pairings; GPT-4o-mini remains low-cost across prompts, and structured prompts (CoT/CoT-few-shot) on GPT-4o-mini offer attractive F1 at a small incremental cost. We recommend a staged workflow that (1) deploys low-cost models with structured prompts for first-pass screening and (2) escalates only borderline cases to higher-capacity models. These findings highlight LLMs' uneven but promising potential to automate literature screening. By systematically analyzing prompt-model interactions, we provide a comparative benchmark and practical guidance for task-adaptive LLM deployment.

文献筛选大模型应用提示工程自动化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。