用AI检测论文评审是否机器生成,防止创新被算法扼杀
Are we still able to recognize pearls? Machine-driven peer review and the risk to creativity: An explainable RAG-XAI detection framework with markers extraction
- 构建RAG-XAI框架,通过提取特征标记识别机器生成评审
- 模型准确率达99.61%,误报率低于0.23%,远超传统方法
- 适合关注科研公平性与AI伦理的研究者参考
大型语言模型(LLMs)融入同行评审引发新担忧:评审自动化可能延伸至整个编辑流程,导致决策权转移至算法系统,重塑科学评价标准。本文指出,机器评估可能系统性偏好标准化、模式化研究,打压非主流和范式突破性思想,诱发知识同质化。为此,提出可解释的RAG-XAI检测框架,通过标记提取保留科学透明性与创造力。该框架在测试集上表现优异:XGBoost、随机森林和LightGBM模型准确率99.61%,AUC-ROC高于0.999,F1分数达0.9925,误报率<0.23%,漏报率约0.8%;而逻辑回归基线仅89.97%准确率,F1为0.8314。特征重要性与SHAP分析显示,个人信号缺失与重复模式是主要判别特征。RAG模块达到90.5%的top-1检索准确率,嵌入空间中同类样本聚类紧密,验证了框架可靠性。
原文摘要 · Abstract (English)
The integration of large language models (LLMs) into peer review raises a concern beyond authorship and detection: the potential cascading automation of the entire editorial process. As reviews become partially or fully machine-generated, it becomes plausible that editorial decisions may also be delegated to algorithmic systems, leading to a fully automated evaluation pipeline. They risk reshaping the criteria by which scientific work is assessed. This paper argues that machine-driven assessment may systematically favor standardized, pattern-conforming research while penalizing unconventional and paradigm-shifting ideas that require contextual human judgment. We consider that this shift could lead to epistemic homogenization, where researchers are implicitly incentivized to optimize their work for algorithmic approval rather than genuine discovery. To address this risk, we introduce an explainable framework (RAG-XAI) for assessing review quality and detecting automated patterns using markers LLM extractor, aiming to preserve transparency, accountability and creativity in science. The proposed framework achieves near-perfect detection performance, with XGBoost, Random Forest and LightGBM reaching 99.61% accuracy, AUC-ROC above 0.999 and F1-scores of 0.9925 on the test set, while maintaining extremely low false positive rates (<0.23%) and false negative rates (~0.8%). In contrast, the logistic regression baseline performs substantially worse (89.97% accuracy, F1-score 0.8314). Feature importance and SHAP analyses identify absence of personal signals and repetition patterns as the dominant predictors. Additionally, the RAG component achieves 90.5% top-1 retrieval accuracy, with strong same-class clustering in the embedding space, further supporting the reliability of the framework's outputs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。