用大模型评估临床试验报告质量,准确率达85%
Evaluation of Clinical Trials Reporting Quality using Large Language Models
- 基于CONSORT标准构建评估语料库,测试大模型判断报告质量
- 最佳模型+提示法组合达85%准确率,链式思考提升推理可解释性
- 适合医学信息学、AI辅助科研与临床研究质控人员参考
报告质量是临床试验研究文章的重要议题,直接影响临床决策。本文测试大语言模型在使用 CONSORT 标准评估此类文章报告质量方面的能力。我们从两项关于摘要报告质量的研究中构建了 CONSORT-QA 评估语料库,采用 CONSORT-abstract 标准。随后评估不同通用领域或生物医学领域适配的大规模生成语言模型,在使用多种已知提示方法(包括链式思考)时正确判断 CONSORT 条目的能力。最佳模型与提示方法组合达到 85% 的准确率。使用链式思考能够提供模型完成任务时的推理过程信息,增强可解释性。
原文摘要 · Abstract (English)
Reporting quality is an important topic in clinical trial research articles, as it can impact clinical decisions. In this article, we test the ability of large language models to assess the reporting quality of this type of article using the Consolidated Standards of Reporting Trials (CONSORT). We create CONSORT-QA, an evaluation corpus from two studies on abstract reporting quality with CONSORT-abstract standards. We then evaluate the ability of different large generative language models (from the general domain or adapted to the biomedical domain) to correctly assess CONSORT criteria with different known prompting methods, including Chain-of-thought. Our best combination of model and prompting method achieves 85% accuracy. Using Chain-of-thought adds valuable information on the model's reasoning for completing the task.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。