测试大模型当审稿人表现,发现它们会高估弱论文且易受隐藏指令攻击。
LLM-as-a-Reviewer: Benchmarking Their Ability, Divergence, and Prompt Injection Resistance as Paper Reviewers

- 在898篇顶会论文上测试12个大模型,评估评分准确性、与人类差异和抗攻击能力。
- 模型普遍高估弱论文,对可复现性过度打分,且评语更冗长、词汇更单一。
- 隐藏指令可显著提升低分论文评分,不同模型家族效果差异明显,需设安全机制。
大型语言模型(LLMs)正越来越多地用于学术同行评审,但其可靠性、与人类判断的一致性以及对对抗攻击的鲁棒性仍不明确。我们针对来自NeurIPS和ICLR的898篇论文,系统性地评估了12个LLM在三个维度上的表现:评分校准度、与人类审稿人的差异性,以及对通过隐形字体映射攻击嵌入的提示注入的抵抗能力。结果表明,LLMs系统性地高估较弱投稿,在主题关注点上与人类存在分歧,低估清晰度而高估可复现性;同时生成的评审意见比人类长2到3倍,词汇多样性更低,用词更标准化。提示注入攻击依然高度有效:简单的隐藏指令即可使大量低分论文被提升至接受级别,且不同模型家族间效果差异显著。尽管LLMs在结构化评审中具有实用价值,但将其引入同行评审仍需防范内在偏差与对抗风险。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly used in academic peer review, yet their reliability, alignment with human judgment, and robustness to adversarial attacks remain poorly understood. We present a systematic benchmark of LLM-as-a-Reviewer on 898 papers stratified from NeurIPS and ICLR, evaluating 12 LLMs along three axes: rating calibration, divergence from human reviewers, and resistance to prompt injection embedded via an invisible font-mapping attack. We find that LLMs systematically overrate weaker submissions and diverge from humans in topical emphasis, under-flagging Clarity and over-flagging Reproducibility, while producing reviews two to three times longer with lower lexical diversity and a more standardized vocabulary. Prompt injection remains highly effective. Simple hidden instructions can promote low-scoring papers to acceptance-level ratings in a substantial fraction of cases, with effectiveness varying sharply across model families. While LLMs offer utility in structuring evaluations, their integration into peer review requires safeguards against both intrinsic biases and adversarial risks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。