对比人类与大模型在复杂文献筛选中的表现,发现配置远比模型本身更重要。
Evaluating human and LLM screening workflows in a conceptually complex scoping review: Recall--workload trade-offs and run-to-run consistency
- 用不同模型和处理方式测试大模型筛选,发现工作流配置影响更大。
- 最高召回率83.9%但保留率超一半,人类工作流保留42%-45%且召回82.3%-82.9%。
- 相同模型重复运行结果不一致,提示需人工监督确保可靠性。
大型语言模型(LLMs)在证据综合中的筛选任务中日益普及,但漏筛可能导致相关研究被遗漏。本研究在一项概念复杂的系统评价中,对人类与大模型的标题与摘要筛选流程进行了预注册比较。经过保守的仅标题筛选后,1,131条记录由一名评审负责人、四名训练助手(非重叠子集)及七次完整的大模型运行(使用不同模型与处理配置,包括一次名义上相同的重复运行)进行筛选。比较了保留的工作量、相对于316条已验证合格记录的操作性召回率、一致性、运行间一致性及流程负担。由于仅对父级评审中推进并评估的记录进行了合格性验证,召回率估计为操作性。结果显示,无一工作流恢复所有已验证合格记录。人类工作流与两次GPT-5.4文件批处理运行分别保留42.2%-45.0%的记录,实现82.3%-82.9%的召回率。Gemini 3.1文件批处理达到最高召回率(83.9%),但保留率高达56.7%。一次性处理配置的合格记录回收量低于对应文件批处理配置。两次名义上相同的GPT-5.4文件批处理运行在91.7%的记录上达成一致,但在94条记录上存在分歧,其中包括29条仅被其中一次运行保留的已验证合格记录。讨论指出,大模型筛选表现取决于实施的工作流,而非模型本身。处理配置、工作量、记录级差异及人机决策融合是部署系统的重要属性。对于高召回需求任务,大模型更适合经过验证、可审计的人工监督工作流,而非自主排除。
原文摘要 · Abstract (English)
Background. Large language models (LLMs) are increasingly used for screening in evidence synthesis, where false negatives can remove relevant studies before full-text assessment. We compared human and LLM title-and-abstract screening workflows in a preregistered study embedded in a conceptually complex scoping review. Methods. After a conservative title-only screen, 1,131 records were screened by one review lead, four trained assistants screening non-overlapping subsets, and seven complete LLM runs using different models and processing configurations, including a nominally identical repeat run. We compared retained workload, operational recall against 316 verified eligible records, agreement, run-to-run consistency, and procedural burden. Because eligibility was verified only for records advanced and assessed in the parent review, recall estimates were operational. Results. No workflow recovered all verified eligible records. The human workflows and two GPT-5.4 file-batch runs retained 42.2-45.0% of records while achieving 82.3-82.9% recall. Gemini 3.1 file batches achieved the highest recall (83.9%) but retained 56.7% of records. All-at-once configurations recovered fewer eligible records than corresponding file-batch configurations. Two nominally identical GPT-5.4 file-batch runs agreed on 91.7% of records but differed on 94 records, including 29 verified eligible records retained by only one run. Discussion. LLM screening performance depended on the implemented workflow, not model identity alone. Processing configuration, workload, record-level variation, and human-LLM decision integration are therefore substantive properties of deployed systems. For high-recall tasks, LLMs are better suited to validated, auditable, human-supervised workflows than autonomous exclusion.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。