微调大模型提升文献筛选效率,准确率接近人工水平。
Fine-Tuning A Large Language Model for Systematic Review Screening
- 用12亿参数模型在8500条数据上微调,增强筛选上下文理解。
- 筛选准确率比基础模型提升80.79%,与人工一致率达86.4%。
- 适合需要高效处理海量文献的系统综述研究者使用。
传统系统综述耗时耗力,主要因需人工审阅大量标题和摘要。近期研究尝试用大语言模型(LLM)提升效率,但结果不一。我们认为原因在于仅靠提示(prompting)无法提供足够上下文。本研究针对系统综述中的研究筛选任务,对一个12亿参数的开源大模型进行微调,基于人类对超过8500条标题和摘要的纳入评分。结果显示,微调后模型性能显著提升,加权F1分数较基线模型提高80.79%。在全部8,277项研究上测试,模型与人工编码者的一致率达到86.40%,真阳性率为91.18%,真阴性率为86.38%,多次推理结果完全一致。结果表明,为大规模系统综述微调大模型具有明显潜力。
原文摘要 · Abstract (English)
Systematic reviews traditionally have taken considerable amounts of human time and energy to complete, in part due to the extensive number of titles and abstracts that must be reviewed for potential inclusion. Recently, researchers have begun to explore how to use large language models (LLMs) to make this process more efficient. However, research to date has shown inconsistent results. We posit this is because prompting alone may not provide sufficient context for the model(s) to perform well. In this study, we fine-tune a small 1.2 billion parameter open-weight LLM specifically for study screening in the context of a systematic review in which humans rated more than 8500 titles and abstracts for potential inclusion. Our results showed strong performance improvements from the fine-tuned model, with the weighted F1 score improving 80.79% compared to the base model. When run on the full dataset of 8,277 studies, the fine-tuned model had 86.40% agreement with the human coder, a 91.18% true positive rate, a 86.38% true negative rate, and perfect agreement across multiple inference runs. Taken together, our results show that there is promise for fine-tuning LLMs for title and abstract screening in large-scale systematic reviews.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。