arXiv:2504.11042cs.CL2025-04中稿 · ACL被引 17

构建首个中文可读的懒惰审稿行为数据集,助力提升论文评审质量。

LazyReview A Dataset for Uncovering Lazy Thinking in NLP Peer Reviews

  • 构建细粒度标注的懒惰审稿句子数据集
  • 指令微调使模型性能提升10-20个百分点
  • 适合训练新手审稿人与改进评审流程

同行评审是科学出版质量控制的核心。随着工作量增加,审稿人无意中使用‘快速启发式’(懒惰思维)的问题日益突出,损害评审质量。现有研究缺乏针对该问题的NLP方法,且无真实世界数据集支持检测工具开发。本文提出LazyReview,一个对同行评审语句进行细粒度懒惰思维类别标注的数据集。分析显示,大语言模型在零样本设置下难以识别此类行为,但通过本数据集进行指令微调后,性能提升10-20个百分点,凸显高质量训练数据的重要性。进一步控制实验表明,接受懒惰思维反馈修订后的评审更具全面性与可操作性。我们将公开数据集及增强版指南,以支持社区培训初级审稿人。(代码见:https://github.com/UKPLab/acl2025-lazy-review)

原文摘要 · Abstract (English)

Peer review is a cornerstone of quality control in scientific publishing. With the increasing workload, the unintended use of `quick' heuristics, referred to as lazy thinking, has emerged as a recurring issue compromising review quality. Automated methods to detect such heuristics can help improve the peer-reviewing process. However, there is limited NLP research on this issue, and no real-world dataset exists to support the development of detection tools. This work introduces LazyReview, a dataset of peer-review sentences annotated with fine-grained lazy thinking categories. Our analysis reveals that Large Language Models (LLMs) struggle to detect these instances in a zero-shot setting. However, instruction-based fine-tuning on our dataset significantly boosts performance by 10-20 performance points, highlighting the importance of high-quality training data. Furthermore, a controlled experiment demonstrates that reviews revised with lazy thinking feedback are more comprehensive and actionable than those written without such feedback. We will release our dataset and the enhanced guidelines that can be used to train junior reviewers in the community. (Code available here: https://github.com/UKPLab/acl2025-lazy-review)

同行评审懒惰思维数据集LLM训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。