测试大模型当审稿人,发现评分偏高且错漏难辨。
Do large language models scrutinise what they review? A multimodal audit of scoring calibration, error detection, and author-identity effects

- 用文本+图文评估大模型审稿能力,对比真实人类评分
- 模型平均分7.0-8.1,远高于人类3.4-6.8,误检率超七成
- 加一句提示可提错检率,但图像对判断无帮助
大型语言模型(LLMs)被越来越多用于生成同行评审意见,引发对其批判性评估能力的关注。本研究评估了两个多模态LLM——Qwen2.5-VL-72B和Pixtral-Large-124B——作为审稿人,在165篇提交至2026年国际学习表征会议(ICLR 2026)的论文上的表现。这些论文的发布日期晚于两模型训练截止时间。论文以盲审、高声誉机构署名或低声誉机构署名形式呈现,格式为纯文本或带图文本。同时,在55篇论文中插入145个可验证错误,以评估在自然提示与验证导向提示下的错误识别能力。所有稿件组中,模型评分范围为7.0–8.1,而人类均分仅为3.4–6.8。自然提示下,模型仅检测到12.1%的错误;增加一句验证提示后提升至22.2%,仍有78%未被发现。提供图表反而降低错误识别率并提高评分。没有视觉错误能可靠对应其图示,一半纯文本评审描述了不存在的图。作者身份不影响评分或错误检测。模型的编辑决策与简单平均得分一致。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly used to generate peer reviews, prompting examination of their capacity for critical evaluation. This study evaluates two multimodal LLMs, Qwen2.5-VL-72B and Pixtral-Large-124B, as reviewers across 165 submissions to the 2026 International Conference on Learning Representations, a venue that postdates both models' training cutoffs. Manuscripts were presented to both models with author identities blinded, replaced with high-prestige affiliations, or replaced with low-prestige affiliations, and in either text-only or text-with-figure format. Additionally, 145 verifiably detectable errors were inserted into 55 manuscripts to assess error identification under natural and verification-oriented prompts. Across all manuscript groups, including rejected submissions, LLM scores ranged from 7.0 to 8.1, whereas human mean scores ranged from 3.4 to 6.8. The models detected 12.1\% of the verified errors under natural prompting, and a one-sentence verification instruction increased detection to 22.2\%; however, 78\% of the errors remained undetected. Providing figures reduced error detection while increasing review scores. No visual error was reliably verified against its corresponding figure, and half of the text-only reviews described figures that were not provided. Author identity did not influence either review scores or error detection. LLM editorial decisions exactly matched those produced by simple score averaging.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。