arXiv:2602.10118cs.CLcs.CY2026-02综述

用大模型分析审稿意见,自动发现并改进问题,提升审稿质量。

Reviewing the Reviewer: Elevating Peer Review Quality through LLM-Guided Feedback

  • 拆解审稿内容为论证片段,结合大模型与传统分类器识别多类问题
  • 生成针对性改进建议,实验显示审稿质量最高提升92.4%
  • 适用于希望提高审稿水平的研究者和期刊编辑

同行评审是科学质量的核心,但依赖简单启发式方法——懒惰思维——已导致标准下降。以往工作将懒惰思维检测视为单标签任务,但审稿段落可能同时存在清晰度、具体性等多重问题。将检测转化为可操作的改进需具备指南意识的反馈,目前尚缺。我们提出一种基于大模型的框架:将审稿意见分解为论证片段,通过神经符号模块结合大模型特征与传统分类器识别问题,并利用遗传算法优化的问题专属模板生成精准反馈。实验表明,该方法优于零样本大模型基线,审稿质量最高提升92.4%。我们还发布了LazyReviewPlus数据集,包含1,309条标注了懒惰思维与具体性的句子。

原文摘要 · Abstract (English)

Peer review is central to scientific quality, yet reliance on simple heuristics -- lazy thinking -- has lowered standards. Prior work treats lazy thinking detection as a single-label task, but review segments may exhibit multiple issues, including broader clarity problems, or specificity issues. Turning detection into actionable improvements requires guideline-aware feedback, which is currently missing. We introduce an LLM-driven framework that decomposes reviews into argumentative segments, identifies issues via a neurosymbolic module combining LLM features with traditional classifiers, and generates targeted feedback using issue-specific templates refined by a genetic algorithm. Experiments show our method outperforms zero-shot LLM baselines and improves review quality by up to 92.4\%. We also release LazyReviewPlus, a dataset of 1,309 sentences labeled for lazy thinking and specificity.

审稿优化大模型应用神经符号

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。