arXiv:2510.11409cs.LGcs.DL2025-10被引 8

用大模型+人工协作,高效筛选文献,降低错误率。

Leveraging LLMs for Semi-Automatic Corpus Filtration in Systematic Literature Reviews

  • 多模型联合判断,通过共识机制筛选论文
  • 8000+篇文献测试,错误率低于单人人工
  • 开源工具支持实时干预,适合科研人员使用

系统性文献综述(SLR)对把握研究领域格局和指引未来方向至关重要,但其文献检索与筛选过程高度耗时且依赖大量人工。本文提出一种基于多个大语言模型(LLMs)的流水线方法,通过描述性提示对论文进行分类,并采用共识机制联合决策。整个流程由人工监督,通过开源可视化分析网页工具LLMSurver实现交互式控制,支持对模型输出的实时检查与修改。我们在包含超过8000篇候选论文的真实SLR数据集上评估该方法,对比了2024年中至2025年秋的主流开源与商用LLM。结果表明,该方法显著减少人工工作量,同时错误率低于单一人类标注者。此外,现代开源模型已足以胜任此任务,使方法具备可及性与成本优势。整体证明,负责任的人机协同能有效加速并提升学术工作中的系统性文献综述质量。

原文摘要 · Abstract (English)

The creation of systematic literature reviews (SLR) is critical for analyzing the landscape of a research field and guiding future research directions. However, retrieving and filtering the literature corpus for an SLR is highly time-consuming and requires extensive manual effort, as keyword-based searches in digital libraries often return numerous irrelevant publications. In this work, we propose a pipeline leveraging multiple large language models (LLMs), classifying papers based on descriptive prompts and deciding jointly using a consensus scheme. The entire process is human-supervised and interactively controlled via our open-source visual analytics web interface, LLMSurver, which enables real-time inspection and modification of model outputs. We evaluate our approach using ground-truth data from a recent SLR comprising over 8,000 candidate papers, benchmarking both open and commercial state-of-the-art LLMs from mid-2024 and fall 2025. Results demonstrate that our pipeline significantly reduces manual effort while achieving lower error rates than single human annotators. Furthermore, modern open-source models prove sufficient for this task, making the method accessible and cost-effective. Overall, our work demonstrates how responsible human-AI collaboration can accelerate and enhance systematic literature reviews within academic workflows.

文献筛选大模型应用人机协同

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。