arXiv:2503.13857cs.CL2025-03被引 3

用AI自动评估预印本质量,提升系统综述的时效性和覆盖面。

Enabling Inclusive Systematic Reviews: Incorporating Preprint Articles with Large Language Model-Driven Evaluations

  • 通过NLP和大模型自动提取预印本文档特征,减少人工标注。
  • 结合语义嵌入与大模型评分,预测准确率达0.747(AUROC)。
  • 适合需要快速整合高质量预印本的研究者使用。

背景:比较有效性研究中的系统综述需要及时整合证据,而预印本虽加速知识传播,但质量参差不齐,给系统综述带来挑战。方法:提出AutoConfidence框架,用于自动化预测预印本发表可能性,减少对人工筛选的依赖,并引入三项改进:(1) 使用自然语言处理技术实现自动化数据提取;(2) 标题与摘要的语义嵌入;(3) 大语言模型驱动的评估得分。同时采用两种预测模型:随机森林分类器(二分类)与生存治愈模型(预测二分类结果及随时间变化的发表风险)。结果:随机森林分类器在使用大模型评分时达到AUROC 0.692,加入语义嵌入后提升至0.733,再加入文章使用指标后达0.747;生存治愈模型在大模型评分下取得AUROC 0.716,加入语义嵌入后升至0.731;在发表风险预测方面,一致性指数为0.658,提升至0.667。结论:通过自动化数据提取与多特征融合,AutoConfidence显著提升了预印本发表预测性能,降低了人工标注负担。该框架有助于在系统综述的评估阶段更有效地纳入预印本资源,支持研究者高效利用预印本知识。

原文摘要 · Abstract (English)

Background. Systematic reviews in comparative effectiveness research require timely evidence synthesis. Preprints accelerate knowledge dissemination but vary in quality, posing challenges for systematic reviews. Methods. We propose AutoConfidence (automated confidence assessment), an advanced framework for predicting preprint publication, which reduces reliance on manual curation and expands the range of predictors, including three key advancements: (1) automated data extraction using natural language processing techniques, (2) semantic embeddings of titles and abstracts, and (3) large language model (LLM)-driven evaluation scores. Additionally, we employed two prediction models: a random forest classifier for binary outcome and a survival cure model that predicts both binary outcome and publication risk over time. Results. The random forest classifier achieved AUROC 0.692 with LLM-driven scores, improving to 0.733 with semantic embeddings and 0.747 with article usage metrics. The survival cure model reached AUROC 0.716 with LLM-driven scores, improving to 0.731 with semantic embeddings. For publication risk prediction, it achieved a concordance index of 0.658, increasing to 0.667 with semantic embeddings. Conclusion. Our study advances the framework for preprint publication prediction through automated data extraction and multiple feature integration. By combining semantic embeddings with LLM-driven evaluations, AutoConfidence enhances predictive performance while reducing manual annotation burden. The framework has the potential to facilitate incorporation of preprint articles during the appraisal phase of systematic reviews, supporting researchers in more effective utilization of preprint resources.

系统综述预印本大模型自动化评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。