分析大模型在文献筛选中的失败原因,提出可落地的改进方案。
Understanding LLMs in Title-Abstract Screening: From Disagreements to Recommendations
- 通过六项软件工程综述研究,对比人类与大模型筛选结果
- 发现大模型在关键词理解、边界判断上存在系统性偏差
- 建议部署前验证语义理解能力,重点处理模糊案例
多项研究探讨了大语言模型(LLMs)在系统综述(SRs)标题摘要筛选中的应用,但准确性参差不齐。本研究超越量化一致率指标,从定性角度分析大模型失败的原因。我们分析了六个软件工程领域系统综述中超过1000篇原始论文的人类专家与大模型零样本筛选结果,其一致性Kappa值介于0.52至0.77之间。定性分析表明,人机分歧主要源于可识别的共性问题:关键术语边界模糊、关键词过度强调及主题误判。基于此,我们提出具体建议:部署前验证语义理解能力,采用多模型并行筛选,将验证重点放在边界案例。未来研究需进一步验证建议效果,行业需共同制定大模型在系统综述中的使用规范。
原文摘要 · Abstract (English)
Several studies have examined the use of large language models (LLMs) for title-abstract screening in systematic reviews (SRs), reporting mixed accuracy. However, questions of reliability remain largely unaddressed. In this study, we go beyond quantitative LLM-human agreement metrics and qualitatively investigate how and why LLMs fail. We also propose actionable recommendations. We analyzed disagreements between LLMs and researchers across six software engineering SRs and over 1,000 primary study papers. For each SR, papers were screened independently by human experts and LLMs in zero-shot mode, resulting in Kappa values ranging from 0.52 to 0.77. Qualitative analysis suggests that human-LLM disagreement results from recurring, identifiable causes, such as boundary ambiguity in key terms, keyword overemphasization, and incorrect topic inference. Based on these findings, we propose recommendations such as validating semantic understanding before deployment, running multiple LLMs, and focusing validation efforts on borderline cases. Future studies are needed to validate the impact of our recommendations, and community efforts are needed to develop normative guidelines on LLM usage in SRs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。