用大模型评估风湿病研究的报告规范,发现它能搞定基础检查但复杂问题靠不住。
Agreement Between Large Language Models, Human Reviewers, and Authors in Evaluating STROBE Checklists for Observational Studies in Rheumatology
- 让大模型、专家和作者各自独立打分,对比对22项STROBE条目的判断一致性。
- 整体一致率达85%,但在方法学复杂条目上大模型与专家差异明显,最低一致系数-0.252。
- 适合用于快速筛查格式化内容,但不能替代专业人员对研究设计的深度评判。
背景:评估观察性研究是否符合流行病学报告强化声明(STROBE)标准耗时且主观。本研究比较了大型语言模型(LLMs)、人类评审团及原始作者在风湿病观察性研究中的STROBE评估结果。方法:基于GRRAS与DEAL Pathway B框架,17篇风湿病文章由作者、五人制人类评审团(从初级到高级专业人员)以及两个大模型(ChatGPT-5.2、Gemini-3Pro)独立评估。使用22项STROBE检查清单,分为方法学严谨性与呈现与背景两域。采用Gwet一致性系数(AC1)计算评分者间一致性。结果:所有评审者总体一致性为85.0%(AC1=0.826)。分域分析显示,呈现与背景域几乎完全一致(AC1=0.841),方法学严谨性域为高度一致(AC1=0.803)。尽管大模型在标准格式条目上与所有人类评审者达成完全一致(AC1=1.000),但在复杂条目上一致性下降。例如,关于失访情况的条目,Gemini 3 Pro与资深评审者的一致性系数为AC1=-0.252,与作者仅达中等水平。此外,ChatGPT-5.2在特定方法学条目上普遍比Gemini-3Pro更接近人类评审者。结论:大模型在基础筛查方面具有潜力,但在复杂方法学条目的表现不佳,可能源于对表面信息的依赖。目前其更适合标准化简单核查,而非取代专家对观察性研究的评价。
原文摘要 · Abstract (English)
Introduction: Evaluating compliance with the Strengthening the Reporting of Observational Studies in Epidemiology (STROBE) statement can be time-consuming and subjective. This study compares STROBE assessments from large language models (LLMs), a human reviewer panel, and the original manuscript authors in observational rheumatology research. Methods: Guided by the GRRAS and DEAL Pathway B frameworks, 17 rheumatology articles were independently assessed. Evaluations used the 22-item STROBE checklist, completed by the authors, a five-person human panel (ranging from junior to senior professionals), and two LLMs (ChatGPT-5.2, Gemini-3Pro). Items were grouped into Methodological Rigor and Presentation and Context domains. Inter-rater reliability was calculated using Gwet's Agreement Coefficient (AC1). Results: Overall agreement across all reviewers was 85.0% (AC1=0.826). Domain stratification showed almost perfect agreement for Presentation and Context (AC1=0.841) and substantial agreement for Methodological Rigor (AC1=0.803). Although LLMs achieved complete agreement (AC1=1.000) with all human reviewers on standard formatting elements, their agreement with human reviewers and authors declined on complex items. For example, regarding the item on loss to follow-up, the agreement between Gemini 3 Pro and the senior reviewer was AC1=-0.252, while the agreement with the authors was only fair. Additionally, ChatGPT-5.2 generally demonstrated higher agreement with human reviewers than Gemini-3Pro on specific methodological items. Conclusion: While LLMs show potential for basic STROBE screening, their lower agreement with human experts on complex methodological items likely reflects a reliance on surface-level information. Currently, these models appear more reliable for standardizing straightforward checks than for replacing expert human judgment in evaluating observational research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。