arXiv:2511.03048cs.CL2025-11EMNLP被引 1

用大模型辅助评估临床试验偏倚风险,提升效率与准确性。

ROBoto2: An Interactive System and Dataset for LLM-assisted Clinical Trial Risk of Bias Assessment

  • 构建交互式网页平台,结合PDF解析与LLM推理自动回答偏倚问题。
  • 发布包含521篇儿科试验、8954个问题的标注数据集,支持模型评测。
  • 适合医学研究者与AI开发者,推动系统评价自动化发展。

我们提出ROBOTO2,一个开源的基于Web的大语言模型(LLM)辅助临床试验偏倚风险(ROB)评估平台。该平台通过交互式界面,整合PDF解析、检索增强的LLM提示和人机协作审核,简化了传统的ROB v2(ROB2)标注流程。用户可上传临床试验报告,获得针对ROB2信号问题的初步答案及支持证据,并实时反馈或修正系统建议。ROBOTO2已公开部署于https://roboto2.vercel.app/,代码与数据同步发布,促进可复现性与应用推广。我们构建并发布了包含521篇儿科临床试验报告的数据集(共8954个信号问题,1202段证据文本),采用人工与LLM协同标注方式,作为基准数据集以支持未来研究。基于该数据集,我们对4种主流LLM在ROB2任务上的表现进行了评测,并分析当前模型能力与自动化系统评价中的关键挑战。

原文摘要 · Abstract (English)

We present ROBOTO2, an open-source, web-based platform for large language model (LLM)-assisted risk of bias (ROB) assessment of clinical trials. ROBOTO2 streamlines the traditionally labor-intensive ROB v2 (ROB2) annotation process via an interactive interface that combines PDF parsing, retrieval-augmented LLM prompting, and human-in-the-loop review. Users can upload clinical trial reports, receive preliminary answers and supporting evidence for ROB2 signaling questions, and provide real-time feedback or corrections to system suggestions. ROBOTO2 is publicly available at https://roboto2.vercel.app/, with code and data released to foster reproducibility and adoption. We construct and release a dataset of 521 pediatric clinical trial reports (8954 signaling questions with 1202 evidence passages), annotated using both manually and LLM-assisted methods, serving as a benchmark and enabling future research. Using this dataset, we benchmark ROB2 performance for 4 LLMs and provide an analysis into current model capabilities and ongoing challenges in automating this critical aspect of systematic review.

临床试验大模型偏倚评估医学AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。