arXiv:2608.14551cs.CLcs.DL2026-08综述

用辅助模型提升大模型文献筛选的可信度,显著提高效率且不丢召回。

Auxiliary uncertainty signals for LLM-assisted systematic review screening: a benchmark across eight Cohen drug-class reviews

论文配图:Auxiliary uncertainty signals for LLM-assisted systematic review screening: a benchmark across eight Cohen drug-class reviews
图 1 · 摘自论文原文
  • 引入BERT+GCN模型生成结构化不确定性信号,指导大模型筛选决策。
  • 仅对不确定样本二次筛选可实现92%召回率,成本仅基础方案的1.05倍。
  • 实证发现大模型无法自我判断该不该重审,适合需要高可靠性的综述场景。

大型语言模型(LLMs)在系统综述的题摘要筛选中应用日益广泛,但其决策缺乏校准的置信度。本文表明,一个辅助的BERT+GCN分类器能提供结构化的不确定性信号,显著提升LLM筛选效率,并识别出最优提示传递策略。我们在八组来自Cohen(2006)基准的药物类别数据集上,采用3次种子×5折分层交叉验证(共600个折叠级结果)评估了五种不同提示传递条件。该方法通过两个谱检验(代数根和分类悖论)将每篇论文分类为包含、排除或可能。条件包括信息量(无/标签/完整分数)、选择性(全部样本或仅可能)和时机(主动或被动两阶段)。跨模型初步实验对比gpt-4.1-mini在三个数据集上的表现。主要发现:(i) 完整上下文传递在F1提升0.011(配对威尔科克森检验p=0.008)、WSS@95提升0.050(p=0.039)的同时,仅增加1.28倍令牌成本,且保持高召回率;(ii) 仅对‘可能’样本路由是帕累托最优:平均召回率达0.92,AUC-ROC达0.54,成本仅为基线的1.05倍——仅为完整上下文开销的六分之一;(iii) 两阶段设计使22.2%±8.8%的记录被重新审视,但未改变任何决策(所有数据集与折叠下翻转率为0),证明当前指令微调的LLM无法自我筛选。跨模型实验显示两类大模型均获得0.8%的召回率提升。20,796篇论文的逐篇消融分析表明,双悖论测试可退化为单行逻辑差判据。我们开源完整流程;利用缓存的LLM响应,600次运行可在一小时内复现。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly used for title-abstract screening in systematic reviews, but their decisions lack calibrated uncertainty. We show that an auxiliary BERT+GCN classifier supplies a structured uncertainty signal that improves LLM screening efficiency, and we identify the prompt-delivery strategy that maximises the benefit-to-cost ratio. We evaluate five LLM prompt-delivery conditions on eight drug-class datasets from the Cohen (2006) benchmark using 3 seeds x 5-fold stratified cross-validation (600 fold-level results). A BERT+GCN model trained per fold classifies each test paper as INCLUDE, EXCLUDE, or MAYBE via two spectral tests (algebraic radical and categorical paradox). Conditions vary information content (none / label / full scores), selectivity (all papers vs. MAYBE only), and timing (proactive vs. reactive two-pass). A cross-model pilot against gpt-4.1-mini on three datasets tests cross-generation transfer. Three findings: (i) Full-context delivery yields significant gains in F1 (+0.011, paired Wilcoxon p=0.008) and WSS@95 (+0.050, p=0.039) at a 1.28x token-cost premium, while preserving recall. (ii) MAYBE-only routing is Pareto-optimal: highest mean recall (0.92) and AUC-ROC (0.54) at only 1.05x baseline cost -- one sixth of full-context overhead. (iii) The two-pass design escalates 22.2% +/- 8.8% of records yet never revises its decision (0% flip rate across all datasets and folds), giving decisive evidence that current instruction-tuned LLMs cannot self-triage. The cross-model pilot shows an identical +0.8% recall uplift for both LLM generations. A per-paper ablation across 20,796 observations shows the dual paradox test reduces empirically to a one-line logit-gap criterion. We release the full pipeline; the 600-run experiment replays in under one hour from cached LLM responses.

大模型筛选不确定性建模系统综述效率优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。