大模型识别临床试验报告合规性能力有限,仅能初步筛查良好报告。
Evaluating the Ability of Large Language Models to Identify Adherence to CONSORT Reporting Guidelines in Randomized Controlled Trials: A Methodological Evaluation Study
- 零样本评估150项随机对照试验,用大模型判断是否符合CONSORT标准
- 顶尖模型最高F1仅0.634,对遗漏和不适用项识别率不足0.400
- 适合做初筛工具,但无法替代人工评估报告缺陷
《强化报告试验声明》(CONSORT)是随机对照试验透明高质量报告的全球标准。人工验证其合规性耗时费力,成为同行评审和证据整合的瓶颈。本研究系统评估了当前大型语言模型在零样本条件下识别已发表随机对照试验对2010版CONSORT声明遵循情况的准确性与可靠性。构建了涵盖多个医学领域的150项随机对照试验黄金标准数据集。主要结局为三分类任务的宏平均F1分数,辅以逐项性能指标和定性错误分析。总体模型表现一般。表现最佳的Gemini-2.5-Flash和DeepSeek-R1模型分别获得0.634的宏平均F1分数和0.280、0.282的科恩κ系数,表明与专家共识仅存在中等一致性。类间表现差异显著:多数模型对符合项识别准确率高(F1 > 0.850),但对不符合项和不适用项识别率极低,F1极少超过0.400。值得注意的是,部分知名模型如GPT-4o表现不佳,宏平均F1仅为0.521。大模型在初步筛查合规项目方面具有潜力,但其当前无法可靠检测报告缺失或方法学缺陷,尚不足以替代人类专家进行试验质量的严格评估。
原文摘要 · Abstract (English)
The Consolidated Standards of Reporting Trials statement is the global benchmark for transparent and high-quality reporting of randomized controlled trials. Manual verification of CONSORT adherence is a laborious, time-intensive process that constitutes a significant bottleneck in peer review and evidence synthesis. This study aimed to systematically evaluate the accuracy and reliability of contemporary LLMs in identifying the adherence of published RCTs to the CONSORT 2010 statement under a zero-shot setting. We constructed a golden standard dataset of 150 published RCTs spanning diverse medical specialties. The primary outcome was the macro-averaged F1-score for the three-class classification task, supplemented by item-wise performance metrics and qualitative error analysis. Overall model performance was modest. The top-performing models, Gemini-2.5-Flash and DeepSeek-R1, achieved nearly identical macro F1 scores of 0.634 and Cohen's Kappa coefficients of 0.280 and 0.282, respectively, indicating only fair agreement with expert consensus. A striking performance disparity was observed across classes: while most models could identify compliant items with high accuracy (F1 score > 0.850), they struggled profoundly with identifying non-compliant and not applicable items, where F1 scores rarely exceeded 0.400. Notably, some high-profile models like GPT-4o underperformed, achieving a macro F1-score of only 0.521. LLMs show potential as preliminary screening assistants for CONSORT checks, capably identifying well-reported items. However, their current inability to reliably detect reporting omissions or methodological flaws makes them unsuitable for replacing human expertise in the critical appraisal of trial quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。