用模糊逻辑和大模型提升文献筛选效率,保留证据不足的文档。
An Auditable Pipeline for Fuzzy Full-Text Screening in Systematic Reviews: Integrating Contrastive Semantic Highlighting and LLM Judgment
- 将筛选任务转为模糊决策,通过对比语义和动态阈值生成分级判断。
- 召回率最高达87.5%,全条件通过率50%,较传统方法翻倍提升。
- 适合需要高召回与可审计性的系统综述研究者使用。
文献全文筛选是系统综述的主要瓶颈,因关键证据分散于长篇异构文档中,难以依赖静态二元规则。本文提出一种可审计、可扩展的管道,将纳入/排除决策重构为模糊问题,在非传染性疾病人群健康建模共识报告网络(POPCORN)背景下进行基准测试。文章将文本分割为重叠块并使用领域自适应模型嵌入;对每项标准(人群、干预、结局、研究方法),计算包含-排除对比相似度与模糊度边界,由Mamdani模糊控制器在多标签设置下映射出动态阈值的分级纳入度。大语言模型(LLM)裁判对突出片段进行三级标注、置信度评分及依据标准的推理说明;当证据不足时,仅降低隶属度而非直接排除。在包含16篇全文、3208个文本块的全正黄金集上,模糊系统召回率达81.3%(人群)、87.5%(干预)、87.5%(结局)、75.0%(研究方法),优于统计基线(56.3%-75.0%)和硬阈值基线(43.8%-81.3%)。全标准通过率为50.0%,高于基线的25.0%和12.5%。跨模型理由一致性达98.3%,人机一致率96.1%,试点综述显示91%的评审者间一致性(kappa=0.82),单篇文献筛选时间从约20分钟降至1分钟以内,成本显著降低。结果表明,结合对比强调与LLM裁决的模糊逻辑可实现高召回、稳定推理与全程可追溯。
原文摘要 · Abstract (English)
Full-text screening is the major bottleneck of systematic reviews (SRs), as decisive evidence is dispersed across long, heterogeneous documents and rarely admits static, binary rules. We present a scalable, auditable pipeline that reframes inclusion/exclusion as a fuzzy decision problem and benchmark it against statistical and crisp baselines in the context of the Population Health Modelling Consensus Reporting Network for noncommunicable diseases (POPCORN). Articles are parsed into overlapping chunks and embedded with a domain-adapted model; for each criterion (Population, Intervention, Outcome, Study Approach), we compute contrastive similarity (inclusion-exclusion cosine) and a vagueness margin, which a Mamdani fuzzy controller maps into graded inclusion degrees with dynamic thresholds in a multi-label setting. A large language model (LLM) judge adjudicates highlighted spans with tertiary labels, confidence scores, and criterion-referenced rationales; when evidence is insufficient, fuzzy membership is attenuated rather than excluded. In a pilot on an all-positive gold set (16 full texts; 3,208 chunks), the fuzzy system achieved recall of 81.3% (Population), 87.5% (Intervention), 87.5% (Outcome), and 75.0% (Study Approach), surpassing statistical (56.3-75.0%) and crisp baselines (43.8-81.3%). Strict "all-criteria" inclusion was reached for 50.0% of articles, compared to 25.0% and 12.5% under the baselines. Cross-model agreement on justifications was 98.3%, human-machine agreement 96.1%, and a pilot review showed 91% inter-rater agreement (kappa = 0.82), with screening time reduced from about 20 minutes to under 1 minute per article at significantly lower cost. These results show that fuzzy logic with contrastive highlighting and LLM adjudication yields high recall, stable rationale, and end-to-end traceability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。