用大模型自动检查系统综述是否符合PRISMA 2020标准,准确率超78%。
Large language models for automated PRISMA 2020 adherence checking
- 输入结构化清单可使模型准确率提升至79.7%。
- 最佳模型在全数据集上达95.1%敏感性,但特异性仅49.3%。
- 适合审稿人快速筛查系统综述合规性,仍需人工复核。
评估系统综述对PRISMA 2020指南的遵循情况仍是同行评审中的负担。为解决共享基准缺失问题,我们构建了一个包含108篇知识共享许可的系统综述的版权友好型基准,并在五种输入格式下评估了十种大语言模型(LLMs)。在开发队列中,提供结构化PRISMA 2020清单(Markdown、JSON、XML或纯文本)的准确率为78.7–79.7%,而仅输入论文原文的准确率仅为45.21%(p < 0.0001),且不同结构化格式间无显著差异(p > 0.9)。各模型准确率在70.6%–82.8%之间,灵敏度-特异度权衡明显,并在独立验证队列中重现。随后,我们选择高灵敏度开源模型Qwen3-Max,扩展至完整数据集(n=120),实现95.1%敏感性和49.3%特异性。结构化清单的提供显著提升了基于LLM的PRISMA评估效果,但人类专家验证仍是编辑决策前的必要环节。
原文摘要 · Abstract (English)
Evaluating adherence to PRISMA 2020 guideline remains a burden in the peer review process. To address the lack of shareable benchmarks, we constructed a copyright-aware benchmark of 108 Creative Commons-licensed systematic reviews and evaluated ten large language models (LLMs) across five input formats. In a development cohort, supplying structured PRISMA 2020 checklists (Markdown, JSON, XML, or plain text) yielded 78.7-79.7% accuracy versus 45.21% for manuscript-only input (p less than 0.0001), with no differences between structured formats (p>0.9). Across models, accuracy ranged from 70.6-82.8% with distinct sensitivity-specificity trade-offs, replicated in an independent validation cohort. We then selected Qwen3-Max (a high-sensitivity open-weight model) and extended evaluation to the full dataset (n=120), achieving 95.1% sensitivity and 49.3% specificity. Structured checklist provision substantially improves LLM-based PRISMA assessment, though human expert verification remains essential before editorial decisions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。