测试ChatGPT用标题摘要预测学术评审结果,发现效果因平台而异。
Evaluating the Predictive Capacity of ChatGPT for Academic Peer Review Outcomes Across Multiple Platforms
- 用30次ChatGPT评分平均预测,仅在ICLR中达中等相关性
- 加入全文后对ICLR相关性提升至0.46,对F1000Research仅微升至0.09
- 不同平台需不同输入,链式思维提示反而降低部分平台表现
尽管先前研究显示大语言模型可部分预测同行评审结果,本文引入两个新场景并采用更稳健的方法——对30次ChatGPT评分取平均。结果显示,仅基于提交标题和摘要,平均评分对F1000Research的评审结果预测能力极弱(Spearman's rho=0.00)。而在SciPost Physics中,对有效性、原创性、重要性维度分别呈现弱正相关(rho=0.25, 0.25, 0.20),清晰度相关性较弱(rho=0.08);对于ICLR论文,相关性为中等(rho=0.38)。若包含全文,ICLR相关性提升至0.46,F1000Research小幅上升至0.09,而SciPost LaTeX文件的各维度相关性变化不一。使用链式思维系统提示后,F1000Research相关性微升至0.10,但ICLR降至0.37,SciPost Physics各维度均下降。总体表明,在特定平台下,ChatGPT可生成弱预审质量评估,但其有效性及最优策略随平台差异显著,输入内容也需适配。
原文摘要 · Abstract (English)
While previous studies have demonstrated that Large Language Models (LLMs) can predict peer review outcomes to some extent, this paper builds on that by introducing two new contexts and employing a more robust method - averaging multiple ChatGPT scores. The findings that averaging 30 ChatGPT predictions, based on reviewer guidelines and using only the submitted titles and abstracts, failed to predict peer review outcomes for F1000Research (Spearman's rho=0.00). However, it produced mostly weak positive correlations with the quality dimensions of SciPost Physics (rho=0.25 for validity, rho=0.25 for originality, rho=0.20 for significance, and rho = 0.08 for clarity) and a moderate positive correlation for papers from the International Conference on Learning Representations (ICLR) (rho=0.38). Including the full text of articles significantly increased the correlation for ICLR (rho=0.46) and slightly improved it for F1000Research (rho=0.09), while it had variable effects on the four quality dimension correlations for SciPost LaTeX files. The use of chain-of-thought system prompts slightly increased the correlation for F1000Research (rho=0.10), marginally reduced it for ICLR (rho=0.37), and further decreased it for SciPost Physics (rho=0.16 for validity, rho=0.18 for originality, rho=0.18 for significance, and rho=0.05 for clarity). Overall, the results suggest that in some contexts, ChatGPT can produce weak pre-publication quality assessments. However, the effectiveness of these assessments and the optimal strategies for employing them vary considerably across different platforms, journals, and conferences. Additionally, the most suitable inputs for ChatGPT appear to differ depending on the platform.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。