分析顶会论文评审中评分与文本内容的一致性,发现高分评审更可能拒稿。
Are the confidence scores of reviewers consistent with the review content? Evidence from top conference proceedings in AI
- 用深度学习识别评论中的模糊表达和关注点
- 高信心评分者更倾向拒绝论文,且文本与评分高度一致
- 适合关注学术评审公平性与人工智能评价机制的研究者
同行评审在学术评价中至关重要。顶级人工智能会议使用评审员信心评分以保障评审可靠性,但现有研究缺乏对文本与评分之间细粒度一致性分析,可能遗漏关键细节。本研究利用深度学习与自然语言处理技术,基于人工智能顶会评审数据,在词、句、方面三个层面评估文本与评分的一致性。通过检测模糊表达(hedge)和关注点,分析报告长度、模糊词汇/句子频率、方面提及情况及情感倾向,评估文本与评分的匹配程度。采用相关性、显著性及回归检验,考察信心评分对论文结果的影响。结果显示,在所有层级上均存在高度一致性,回归分析表明更高的信心评分与论文被拒相关,验证了专家判断的有效性与同行评审的公平性。
原文摘要 · Abstract (English)
Peer review is vital in academia for evaluating research quality. Top AI conferences use reviewer confidence scores to ensure review reliability, but existing studies lack fine-grained analysis of text-score consistency, potentially missing key details. This work assesses consistency at word, sentence, and aspect levels using deep learning and NLP conference review data. We employ deep learning to detect hedge sentences and aspects, then analyze report length, hedge word/sentence frequency, aspect mentions, and sentiment to evaluate text-score alignment. Correlation, significance, and regression tests examine confidence scores' impact on paper outcomes. Results show high text-score consistency across all levels, with regression revealing higher confidence scores correlate with paper rejection, validating expert assessments and peer review fairness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。