用ChatGPT评估医学研究质量,发现其对新论文有效但对顶尖期刊效果差。
Evaluating the quality of published medical research with ChatGPT
- 用ChatGPT评分与专家评分对比,验证其评估能力。
- 在临床医学领域,模型评分与机构平均分相关性为r=0.134(n=9872)。
- 适合用于新发表论文的质量评估,尤其替代引用指标。
评估已发表研究质量对部门、研究人员和求职者评价至关重要。引用指标虽常用,但对新论文无效且准确率低。已有研究表明,ChatGPT可预测研究质量,其评分在各领域与专家评分正相关,常优于引用指标,但在临床医学中例外。本文基于最大规模数据集(n=9872)和更深入分析,发现提交至英国科研卓越框架(REF)2021临床医学单位评估(UoA 1)的论文,其ChatGPT 4o-mini评分与机构平均REF评分呈正相关(r=0.134),理论最大相关性为r=0.226。ChatGPT 4o与3.5 turbo亦表现正相关。在机构层面,平均得分与机构均值相关性更强(r=0.395,n=31)。对100本高发文期刊,其平均ChatGPT评分与REF得分强相关(r=0.495),但与引用率负相关(r=-0.148)。期刊与机构层面的异常表明,ChatGPT在评估顶级医学期刊或直接影响人类健康的科研时效果有限。然而结果仍证明其在临床医学整体评估中的潜力,可替代引用指标用于新研究。
原文摘要 · Abstract (English)
Estimating the quality of published research is important for evaluations of departments, researchers, and job candidates. Citation-based indicators sometimes support these tasks, but do not work for new articles and have low or moderate accuracy. Previous research has shown that ChatGPT can estimate the quality of research articles, with its scores correlating positively with an expert scores proxy in all fields, and often more strongly than citation-based indicators, except for clinical medicine. ChatGPT scores may therefore replace citation-based indicators for some applications. This article investigates the clinical medicine anomaly with the largest dataset yet and a more detailed analysis. The results showed that ChatGPT 4o-mini scores for articles submitted to the UK's Research Excellence Framework (REF) 2021 Unit of Assessment (UoA) 1 Clinical Medicine correlated positively (r=0.134, n=9872) with departmental mean REF scores, against a theoretical maximum correlation of r=0.226. ChatGPT 4o and 3.5 turbo also gave positive correlations. At the departmental level, mean ChatGPT scores correlated more strongly with departmental mean REF scores (r=0.395, n=31). For the 100 journals with the most articles in UoA 1, their mean ChatGPT score correlated strongly with their REF score (r=0.495) but negatively with their citation rate (r=-0.148). Journal and departmental anomalies in these results point to ChatGPT being ineffective at assessing the quality of research in prestigious medical journals or research directly affecting human health, or both. Nevertheless, the results give evidence of ChatGPT's ability to assess research quality overall for Clinical Medicine, where it might replace citation-based indicators for new research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。