现有论文评审中使用AI润色的政策难以执行,因检测器常误判人类与AI协作文本。
Policies Permitting LLM Use for Polishing Peer Reviews Are Currently Not Enforceable
- 构建多层级人机协作评审数据集,测试五种先进检测模型。
- 所有检测器均误判非零比例润色文本为AI生成,存在误伤风险。
- 领域特定信号虽有帮助,但无法满足学术诚信识别的准确率要求。
多个科学会议和期刊已出台政策,禁止评审人使用大语言模型(LLM),仅允许用于润色、改写和语法修正。但这些政策是否可执行?我们构建了模拟不同人机协作程度的评审文本数据集,评估了五种前沿检测器(含两种商用系统)。结果表明,所有检测器均会将非零比例的LLM润色文本错误标记为完全由AI生成,可能造成对学术不端的误判。进一步分析发现,利用论文原文访问权限及科学写作领域限制等评审特有信号,在某些场景下可提升检测效果,但每种方法均有局限,均未达到识别评审中AI使用所需的精度标准。研究提示,当前基于检测器估算的评审中AI使用率可能存在夸大,因混合型人机输出常被误判为纯AI生成。
原文摘要 · Abstract (English)
A number of scientific conferences and journals have recently enacted policies that prohibit LLM usage by peer reviewers, except for polishing, paraphrasing, and grammar correction of otherwise human-written reviews. But, are these policies enforceable? To answer this question, we assemble a dataset of peer reviews simulating multiple levels of human-AI collaboration, and evaluate five state-of-the-art detectors, including two commercial systems. Our analysis shows that all detectors misclassify a non-trivial fraction of LLM-polished reviews as AI-generated, thereby risking false accusations of academic misconduct. We further investigate whether peer-review-specific signals, including access to the paper manuscript and the constrained domain of scientific writing, can be leveraged to improve detection. While incorporating such signals yields measurable gains in some settings, we identify limitations in each approach and find that none meets the accuracy standards required for identifying AI use in peer reviews. Importantly, our results suggest that recent public estimates of AI use in peer reviews through the use of AI-text detectors should be interpreted with caution, as current detectors misclassify mixed reviews (collaborative human-AI outputs) as fully AI generated, potentially overstating the extent of policy violations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。