通过多层级扰动分析,揭示大模型在审稿中易受内容缺陷影响的脆弱性。
Aspect-Guided Multi-Level Perturbation Analysis of Large Language Models in Automated Peer Review
- 按贡献、严谨性等维度对论文、评审和回复进行定向扰动
- 发现模型对拒稿结论敏感,且对不完整回复误判为合理
- 适用于审稿系统开发者与模型鲁棒性研究者
我们提出一种基于要素引导的多层级扰动框架,用于评估大语言模型(LLMs)在自动化同行评审中的鲁棒性。该框架针对论文、评审和回应三个关键环节,在贡献、严谨性、表达、语气和完整性等质量维度上实施扰动。通过分析扰动对 LLM-as-Reviewer 与 LLM-as-Meta-Reviewer 的影响,我们发现:仅删除方法细节或改变评审结论即可引发显著偏差;强烈拒稿意见会显著影响元评审,负面或误导性评审被误判为详尽,而不完备或敌意回应反而可能提高接受率。统计检验表明,这些偏差在多种 Chain-of-Thought 提示策略下仍存在,凸显当前模型缺乏稳健的批判性评估能力。本框架为诊断此类脆弱性提供了实用方法,助力构建更可靠、鲁棒的自动评审系统。
原文摘要 · Abstract (English)
We propose an aspect-guided, multi-level perturbation framework to evaluate the robustness of Large Language Models (LLMs) in automated peer review. Our framework explores perturbations in three key components of the peer review process-papers, reviews, and rebuttals-across several quality aspects, including contribution, soundness, presentation, tone, and completeness. By applying targeted perturbations and examining their effects on both LLM-as-Reviewer and LLM-as-Meta-Reviewer, we investigate how aspect-based manipulations, such as omitting methodological details from papers or altering reviewer conclusions, can introduce significant biases in the review process. We identify several potential vulnerabilities: review conclusions that recommend a strong reject may significantly influence meta-reviews, negative or misleading reviews may be wrongly interpreted as thorough, and incomplete or hostile rebuttals can unexpectedly lead to higher acceptance rates. Statistical tests show that these biases persist under various Chain-of-Thought prompting strategies, highlighting the lack of robust critical evaluation in current LLMs. Our framework offers a practical methodology for diagnosing these vulnerabilities, thereby contributing to the development of more reliable and robust automated reviewing systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。