arXiv:2410.21495cs.CL2024-10被引 6

用AI自动判断临床试验偏倚风险,准确率达83%

RoBIn: A Transformer-Based Model For Risk Of Bias Inference With Machine Reading Comprehension

  • 基于Transformer的双任务模型,从论文中提取证据并判断偏倚风险
  • 在多个场景下表现优于传统方法和大模型,最高AUC达0.83
  • 适合需要快速评估文献质量的研究人员和审稿人

科学出版物在揭示新见解、测试新药和制定医疗政策中至关重要。评估其质量需进行偏倚风险(RoB)评估,传统上由人工完成。本研究构建了一个用于机器阅读理解与偏倚评估的新数据集,并提出RoBIn模型,实现自动化评估。该模型采用双任务策略,从文本中提取证据并据此判断偏倚风险。基于科克伦系统综述数据库(CDSR)作为真实标签,对开放获取的临床试验文献(来自PubMed)进行标注,构建了训练与测试数据集。开发了两种基于Transformer的方法:抽取式(RoBInExt)与生成式(RoBInGen),分别用于有效提取证据与分类偏倚风险。实验表明,最佳的RoBIn变体在多数设置下超越传统机器学习与大语言模型方法,达到0.83的ROC AUC。结论:基于临床试验报告中的证据,RoBIn可实现二分类判断——低偏倚或高/不确定偏倚。RoBInGen与RoBInExt均表现稳健,在多种场景下取得最优结果。

原文摘要 · Abstract (English)

Objective: Scientific publications play a crucial role in uncovering insights, testing novel drugs, and shaping healthcare policies. Accessing the quality of publications requires evaluating their Risk of Bias (RoB), a process typically conducted by human reviewers. In this study, we introduce a new dataset for machine reading comprehension and RoB assessment and present RoBIn (Risk of Bias Inference), an innovative model crafted to automate such evaluation. The model employs a dual-task approach, extracting evidence from a given context and assessing the RoB based on the gathered evidence. Methods: We use data from the Cochrane Database of Systematic Reviews (CDSR) as ground truth to label open-access clinical trial publications from PubMed. This process enabled us to develop training and test datasets specifically for machine reading comprehension and RoB inference. Additionally, we created extractive (RoBInExt) and generative (RoBInGen) Transformer-based approaches to extract relevant evidence and classify the RoB effectively. Results: RoBIn is evaluated across various settings and benchmarked against state-of-the-art methods for RoB inference, including large language models in multiple scenarios. In most cases, the best-performing RoBIn variant surpasses traditional machine learning and LLM-based approaches, achieving an ROC AUC of 0.83. Conclusion: Based on the evidence extracted from clinical trial reports, RoBIn performs a binary classification to decide whether the trial is at a low RoB or a high/unclear RoB. We found that both RoBInGen and RoBInExt are robust and have the best results in many settings.

偏倚评估Transformer临床试验AI辅助

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。