提出新评估方法,让大模型文献筛选更可靠。
LLM4SCREENLIT: Recommendations on Assessing the Performance of Large Language Models for Screening Literature in Systematic Reviews
- 用加权马修斯相关系数(WMCC)融合误判代价,改进评估
- 实测显示传统准确率最高模型漏掉63.3%重要文献
- 建议报告完整混淆矩阵,将未分类项视为需人工处理
大语言模型(LLMs)被越来越多用于系统性综述的文献筛选,但标准的混淆矩阵指标在数据不平衡、误判代价不对称的情况下可能误导结果。本文基于对28篇相关论文的分析,提出一种加权马修斯相关系数(WMCC),结合了MCC的随机校正与非对称误判成本,并在三个软件工程领域的再分析中验证,最大规模涵盖9个LLM × 24项二次研究(34,528篇文章)。结果显示,在29篇论文中仅10%报告MCC,24%报告完整混淆矩阵,且无一篇论文为假阴性定价。在最显著的9,695篇文献研究中,准确率最优的LLM丢失63.3%相关证据,MCC最优仅43.9%,而WMCC最优仅5.8%。敏感性分析表明,保守默认权重为10可行。结论建议:应优先关注丢失证据量,使用成本敏感的WMCC与MCC并列排名;报告须包含完整混淆矩阵,未分类输出应视为需人工审查的阳性;设计应防数据泄露,若目标为指导实践且标签可用,应设非LLM基线。
原文摘要 · Abstract (English)
Context: Large language models (LLMs) are increasingly used to screen literature for systematic reviews (SRs), but the standard confusion-matrix metrics used to evaluate them can mislead under the imbalanced, cost-asymmetric conditions of screening. Objective: We develop and justify LLM4SCREENLIT-practical recommendations for researchers conducting LLM-screening evaluations and for editors and reviewers assessing such studies-differentiated by study type (retrospective benchmarking vs deployment for a specific SR). Method: Using Delgado-Chaves et al. (2025), an 18-LLM benchmark across three biomedical SRs, as a motivating example, we reviewed 28 additional papers and extracted their reported metrics. We propose a Weighted Matthews Correlation Coefficient (WMCC) that integrates MCC's chance-correction with asymmetric misclassification costs, and validated it on three software-engineering (SE) reanalyses, the largest covering 9 LLMs x 24 SE secondary studies (34,528 articles). Results: Across the 29 papers, only 10% reported MCC, only 24% reported full confusion matrices, and none of the five papers claiming workload savings priced false-negative cost. In the largest SE reanalysis, MCC and WMCC disagree on the best LLM in 55% of evaluable studies; in the most striking 9,695-article SE study, the Accuracy-best LLM loses 63.3% of relevant evidence (Lost Evidence), the MCC-best 43.9%, but the WMCC-best only 5.8%. Sensitivity analysis (median crossover at w~=2.7, all <7) supports w=10 as a conservative default. Conclusions: SR-screening evaluations should prioritize Lost Evidence and use cost-sensitive WMCC alongside MCC for ranking. Reporting must include the full confusion matrix and treat unclassifiable outputs as positives requiring human review. Designs should be leakage-aware, with non-LLM baselines when the study aims to inform SR practice and labels are available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。