用ChatGPT评估学术研究社会影响,效果因学科而异
Assessing the societal influence of academic research with ChatGPT: Impact case study evaluations
- 输入案例标题与摘要,结合优化版评估指南,让ChatGPT评分
- 跨34个学科相关性达0.18至0.56,部分院系达0.71
- 可辅助专家评审,但需注意学科差异
学术机构常被评估其研究对社会的影响。例如英国科研卓越框架(REF)要求提交五页的证据型影响案例(ICS)。本研究探讨了ChatGPT能否评估此类社会影响声明,从而辅助人类评审。将6,220份公开的REF2021 ICS内容输入ChatGPT 4o-mini,并结合REF2021评估标准进行比较,分析结果与各院系平均评分的相关性。结果显示,仅输入标题与摘要、并优化评估指南时,相关性最高;所有34个评估单位(UoAs)均呈正相关,系数范围为0.18(经济学与计量经济学)至0.56(心理学、精神病学与神经科学)。在院系层面,相关性更高,体育与运动科学、休闲与旅游领域达0.71。表明基于ChatGPT的评估方法简单可行,可用于支持或校验专家判断,但其价值随学科不同而显著变化。
原文摘要 · Abstract (English)
Academics and departments are sometimes judged by how their research has benefitted society. For example, the UK Research Excellence Framework (REF) assesses Impact Case Studies (ICS), which are five-page evidence-based claims of societal impacts. This study investigates whether ChatGPT can evaluate societal impact claims and therefore potentially support expert human assessors. For this, various parts of 6,220 public ICS from REF2021 were fed to ChatGPT 4o-mini along with the REF2021 evaluation guidelines, comparing the results with published departmental average ICS scores. The results suggest that the optimal strategy for high correlations with expert scores is to input the title and summary of an ICS but not the remaining text, and to modify the original REF guidelines to encourage a stricter evaluation. The scores generated by this approach correlated positively with departmental average scores in all 34 Units of Assessment (UoAs), with values between 0.18 (Economics and Econometrics) and 0.56 (Psychology, Psychiatry and Neuroscience). At the departmental level, the corresponding correlations were higher, reaching 0.71 for Sport and Exercise Sciences, Leisure and Tourism. Thus, ChatGPT-based ICS evaluations are simple and viable to support or cross-check expert judgments, although their value varies substantially between fields.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。