研究发现审稿评分比文字反馈更可靠,因文字常受礼貌原则干扰。
Decoupling Scores and Text: The Politeness Principle in Peer Review

- 用3万篇ICLR论文数据对比评分与文字预测效果
- 评分模型准确率91%,文字模型仅81%且依赖大模型
- 拒稿文本仍偏正向,作者难从文字判断结果
作者常难以解读审稿意见,误读礼貌性表述或困惑于具体低分。为此,我们构建了超过3万篇ICLR 2021-2025年投稿数据集,对比基于数值评分与文本评论的接收预测性能。实验显示显著差距:基于评分的模型达到91%准确率,而基于文本的模型即使使用大语言模型也仅达81%,表明文本信息可靠性远低于评分。进一步分析评分模型误判的9%样本,发现其分数分布具有高尖度与负偏态,说明个别低分在拒稿中起决定性作用,即便平均分接近阈值。从情感视角分析文本模型落后原因,揭示‘礼貌原则’普遍存在:被拒稿件的评论仍含更多正面词,掩盖真实拒稿信号,使作者仅凭文字难以判断结果。
原文摘要 · Abstract (English)
Authors often struggle to interpret peer review feedback, deriving false hope from polite comments or feeling confused by specific low scores. To investigate this, we construct a dataset of over 30,000 ICLR 2021-2025 submissions and compare acceptance prediction performance using numerical scores versus text reviews. Our experiments reveal a significant performance gap: score-based models achieve 91% accuracy, while text-based models reach only 81% even with large language models, indicating that textual information is considerably less reliable. To explain this phenomenon, we first analyze the 9% of samples that score-based models fail to predict, finding their score distributions exhibit high kurtosis and negative skewness, which suggests that individual low scores play a decisive role in rejection even when the average score falls near the borderline. We then examine why text-based accuracy significantly lags behind scores from a review sentiment perspective, revealing the prevalence of the Politeness Principle: reviews of rejected papers still contain more positive than negative sentiment words, masking the true rejection signal and making it difficult for authors to judge outcomes from text alone.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。