arXiv:2604.27006cs.SEcs.AI2026-04被引 1

对比12个大模型在软件工程文献筛选中的表现,发现效果差异大且不可靠。

Beyond Accuracy: LLM Variability in Evidence Screening for Software Engineering SLRs

论文配图:Beyond Accuracy: LLM Variability in Evidence Screening for Software Engineering SLRs
图 1 · 摘自论文原文
  • 测试12个大模型和4种传统分类器,评估其在真实文献综述中的筛选表现
  • 摘要信息至关重要,去掉后性能大幅下降,附加标题关键词无明显提升
  • 大模型间差异显著,不比传统方法更优,适用性需结合实际条件判断

背景:系统性文献综述中的研究筛选成本高、易出错,且错误具有不对称风险,假阴性会破坏有效性。尽管大语言模型(LLMs)迅速普及,但其在筛选阶段的表现,特别是不同模型选择及其与传统模型的对比,仍缺乏充分证据。目标:评估大模型在筛选中的表现与变异性,量化输入元数据(摘要、标题、关键词)的影响,并在统一协议下比较大模型与传统分类器的实际收益。方法:我们分析了来自4个厂商(OpenAI、Google Gemini、Anthropic、Llama)的12个大模型,以及4种经典模型(逻辑回归、支持向量机、随机森林、朴素贝叶斯),在两个真实的系统性文献综述(SLRs)上共518篇论文上的表现。实验设计涵盖三个关键维度:(i) 大模型性能的异质性,(ii) 输入特征构成(摘要、标题、关键词)对大模型性能的影响,(iii) 使用大模型相比传统分类器的实际增益。结果:大模型表现出显著的异质性和残余非确定性,即使在温度为零时亦然。摘要的存在与否起决定性作用:移除摘要会持续降低性能,而仅在摘要基础上增加标题和/或关键词并未带来稳健提升。与传统模型相比,性能差异并不一致,无法支持大模型普遍优越的结论。讨论:大模型的采用应基于可复现性、成本、元数据可用性等运营与治理约束,需通过试点验证并明确报告变异性及输入配置。

原文摘要 · Abstract (English)

Context: Study screening in systematic literature reviews is costly, inconsistency-prone, and risk-asymmetric, since false negatives can compromise validity. Despite rapid uptake of Large Language Models (LLMs), there is limited evidence on how such models behave during the study screening phase, particularly regarding the choice of specific LLMs and their comparison with classical models. Objective: To assess LLM performance and variability in screening, quantify the impact of input metadata (abstract, title, keywords), and compare LLMs with classical classifiers under a shared protocol. Methods: We analyzed 12 LLMs from 4 providers (OpenAI, Google Gemini, Anthropic, Llama) and 4 classical models (Logistic Regression, Support Vector Classification, Random Forest, and Naive Bayes) on 2 real Systematic Literature Reviews (SLRs), totaling 518 papers. The experimental design investigated 3 critical dimensions: (i) LLMs performance variability, (ii) the impact of input feature composition (abstract, title, and keywords) on LLM performance, and (iii) the real gain of using LLMs instead of more traditional classification models. Results: LLMs exhibited substantial heterogeneity and residual non-determinism even at temperature zero. Abstract availability was decisive: removing it consistently degraded performance, while adding title and/or keywords to the abstract yielded no robust gains. Compared to classical models, performance differences were not consistent enough to support generalizable LLM superiority. Discussion: LLM adoption should be justified by operational and governance constraints (reproducibility, cost, metadata availability), supported by pilot validation and explicit reporting of variability and input configuration.

大模型评估文献筛选可复现性软件工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。