arXiv:2609.08475cs.CLcs.DL2026-09

研究发现,评审者对复杂词汇的偏好已随大模型兴起而下降。

Do Reviewers Still Reward Lexical Complexity? A Frozen-Rater Study of Preference Drift in 124K ICLR Reviews

论文配图:Do Reviewers Still Reward Lexical Complexity? A Frozen-Rater Study of Preference Drift in 124K ICLR Reviews
图 1 · 摘自论文原文
  • 用固定模型生成评审对比人类评审,分离出偏好变化
  • 人类对非领域词汇复杂度的评分从正相关转为负相关
  • 适合关注学术评审演变与AI评估偏移的研究者

大型语言模型大幅降低了生成复杂文本的成本,评审者是否仍青睐此类表达,反映的是评价者的变化而非文本本身。当写作特征与评分之间的关联随年份变动时,可能是评审者、投稿内容或两者共同变化所致。本文采用冻结评分者方法:使用同一模型家族和提示,在2025年2月至4月间生成81,850条机器评审,覆盖2018至2025年ICLR投稿。通过比较人类评审与机器评审在不同年份的系数差异,可分离出仅由投稿内容变化带来的影响。基于124,615份人类评审(来自32,638篇投稿)分析发现,人类对非领域词汇复杂度的系数从+0.142降至-0.015,而冻结模型则稳定在+0.080至+0.082之间;三重差分结果为-0.0100(q=0.013),40组随机词表置换测试均集中于零。人类仍奖励句长多样性,但机器从未捕捉到此特征;机器仍按早期标准支付词汇复杂度溢价。所有结论均通过双重错误发现控制与区间排除检验,未通过对抗性再测试的结果亦被报告。评审者放弃了生产成本骤降的信号,符合可操控信号模型预测;一个以历史人类偏好校准的LLM裁判会继承旧标准并逐渐偏离,尽管其与人类总体评分的一致性仍正常。

原文摘要 · Abstract (English)

Large language models have collapsed the cost of producing lexically elaborate prose, and whether peer reviewers still reward it is a question about the evaluator, not about the text. When the association between a writing cue and review scores moves across years, the reviewers may have changed, the submissions may have changed, or both, and a regression of scores on text cannot say which. We separate the two with a frozen rater: 81,850 machine reviews of ICLR submissions from 2018 to 2025, all generated in one February-April 2025 window with one model family and one prompt, so that its year-to-year coefficients track submission composition alone and the human-minus-frozen trend difference identifies reviewer preference drift. On 32,638 submissions with 124,615 human reviews, the human coefficient on non-domain lexical complexity falls from +0.142 to -0.015 while the frozen rater moves from +0.080 to +0.082; the three-way difference-in-differences is -0.0100 (q=0.013), and forty random-wordlist placebos through the same specification centre on zero. Humans still reward sentence-length variability, which the frozen rater never registers, while the frozen rater still pays for lexical complexity at its earlier rate. Every claim is held to a double gate of false-discovery control and interval exclusion, and the findings that failed adversarial re-testing are reported. Reviewers discounted a cue whose production cost collapsed, as models of manipulable signals prescribe; an LLM judge calibrated to historical human preferences inherits the earlier schedule and drifts out of alignment while its agreement with humans on totals stays ordinary.

评审偏见大模型评估偏好演化量化分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。