arXiv:2502.01676cs.CLcs.CY2025-02中稿 · NeurIPS综述

构建首个同行评审毒害检测数据集,评测大模型识别学术攻击性评论的能力。

Benchmark on Peer Review Toxic Detection: A Challenging Task with a New Dataset

  • 按四类标准标注开放评审平台数据,构建首个同行评审毒害检测数据集。
  • GPT-4在详细提示下与人工判断一致性达0.63(原为0.56),信心评分高时表现更优。
  • 适合关注学术伦理、AI辅助审稿或可信内容安全的研究者参考。

同行评审对科学进步至关重要,但毒害性反馈会打击作者并阻碍科研发展。本文探索这一重要但研究不足的领域:检测同行评审中的毒性内容。我们定义了四类评审毒害行为,并从OpenReview平台收集标注数据,由人类专家依据上述标准进行标注。基于该数据集,我们评估多种模型,包括专用毒性检测模型、情感分析模型、多个开源大语言模型(LLMs)及两个闭源模型。实验考察不同提示粒度(从粗略到精细)对模型性能的影响。结果显示,如GPT-4等先进LLM在简单提示下与人工判断一致性较低,但在详细指令下显著提升;且模型置信度是判断其与人工判断一致性的良好指标。例如,GPT-4在仅保留置信度高于95%的预测时,Cohen's Kappa得分从0.56提升至0.63。整体表明,当前大模型在该任务上仍有改进空间。本工作旨在推动构建健康、负责任的学术交流环境。

原文摘要 · Abstract (English)

Peer review is crucial for advancing and improving science through constructive criticism. However, toxic feedback can discourage authors and hinder scientific progress. This work explores an important but underexplored area: detecting toxicity in peer reviews. We first define toxicity in peer reviews across four distinct categories and curate a dataset of peer reviews from the OpenReview platform, annotated by human experts according to these definitions. Leveraging this dataset, we benchmark a variety of models, including a dedicated toxicity detection model, a sentiment analysis model, several open-source large language models (LLMs), and two closed-source LLMs. Our experiments explore the impact of different prompt granularities, from coarse to fine-grained instructions, on model performance. Notably, state-of-the-art LLMs like GPT-4 exhibit low alignment with human judgments under simple prompts but achieve improved alignment with detailed instructions. Moreover, the model's confidence score is a good indicator of better alignment with human judgments. For example, GPT-4 achieves a Cohen's Kappa score of 0.56 with human judgments, which increases to 0.63 when using only predictions with a confidence score higher than 95%. Overall, our dataset and benchmarks underscore the need for continued research to enhance toxicity detection capabilities of LLMs. By addressing this issue, our work aims to contribute to a healthy and responsible environment for constructive academic discourse and scientific collaboration.

毒害检测同行评审大模型评测学术伦理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。