rebuttal能改变评审分数,但初始评审结构限制了变化幅度
Rebuttals Move Peer-Review Scores, but Initial-Review Structure Bounds the Movement
- 用LLM从评审文本推断隐含初评分,量化反驳对分数的影响
- 文本暗示评分高于实际分时,分数上升率高达31.9%;低于则仅8.3%
- 真正有效的交流信号多为反驳失败的特征,而非成功沟通
作者反驳是同行评审中唯一的投稿后反馈窗口,但其对评分的影响难以衡量,因分数变动混合了初评位置、论文共识度、审稿人信心及讨论动态。我们基于ICLR 2024-2025的7.3万条审稿轨迹,利用外部存档的前后评分数据,仅将LLM作为测量工具。Gemini Flash 3.0从去评分的评审文本中预测隐含初评分;文本与实际评分的偏差可预测后续分数变化:当文本显示评分低于实际分时,分数上升率为8.3%;高于时达31.9%。Claude Opus 4.6构建并验证了一个包含44个特征的已解决审稿-作者互动分类体系,其中23个特征在模型与独立年份中均通过博尔赫修正检验。在参与反驳的基准集(n=6,705)中,初始评审结构已能较好预测分数变动(AUC=0.747,最小AUC=0.696),加入已解决互动特征后提升至0.804。反驳可推动分数变化,但可测量的变动受限于初始评审结构,且稳健的交流信号多为反驳失败的表现。
原文摘要 · Abstract (English)
Author rebuttals are the main post-submission window in peer review, but their effect on reviewer scores remains hard to measure because score updates mix rebuttal content with initial score position, paper-level consensus, reviewer confidence, and discussion dynamics. We study ICLR 2024-2025 using 73,000 reviewer trajectories with externally archived pre- and post-rebuttal scores, and use LLMs only as measurement instruments. Gemini Flash 3.0 predicts implied pre-rebuttal scores from score-stripped review text. The resulting text-score offset predicts later movement, with score-increase rates rising from 8.3% when text reads below the assigned score to 31.9% when it reads above. Claude Opus 4.6 induces, and outcome-blinded Gemini Flash 3.0 validates, a 44-feature taxonomy of resolved reviewer-author exchanges, where 23 features replicate across model and held-out year under Bonferroni correction. In the rebuttal-engaged benchmark (n=6,705), initial-review structure already predicts much score movement (AUC=0.747, minimal AUC=0.696), while adding the resolved exchange raises AUC to 0.804. Rebuttals can move scores, but measurable movement is bounded by initial-review structure, and robust exchange signals are mostly rebuttal failure modes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。