研究大模型在系统综述筛选中的类别不平衡问题
Class Imbalance and Batch Effects in LLM-Based Screening for Systematic Reviews

- 对比个体与批量处理,分析类别分布对模型行为的影响
- 批量处理显著改变判断行为,且受类别比例影响
- 预估流行率信息无助于提升效果,需关注决策偏差
本研究以系统综述的文献筛选为应用场景,分析大语言模型在类别不平衡二分类任务中的表现。在五个综述中,比较了个体处理与批量处理、是否使用流行率元数据的差异。结果表明,流行率元数据对性能影响有限,无证据显示其能提升效果。相反,批量处理引发显著的行为变化,且变化程度随类别占比不同而异。整体分析与个体样本分析结果并不总是一致。因此,评估批量处理不仅需考虑成本,还应关注其对决策行为的影响。
原文摘要 · Abstract (English)
This study analyses LLMs in imbalanced binary classification, using study screening in systematic reviews as the application domain. An experiment was conducted in five reviews, comparing individual and batch processing, with and without prevalence metadata. The results indicate a limited influence of the prevalence metadata, with no evidence that it improves performance. In contrast, batch processing produced larger behavioral changes that varied according to the prevalence of the class. The aggregate and item-level analyses did not always coincide. Therefore, batch processing should be evaluated not only in terms of cost, but also in relation to its effects on decision-making behavior.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。