对比arXiv上综述与非综述论文的AI生成内容比例,发现非综述类论文受政策影响更大。
LLM-Generated or Human-Written? Comparing Review and Non-Review Papers on ArXiv
- 用两种高精度检测方法分析近年arXiv论文的AI生成比例。
- 综述类论文中AI生成内容占比更高,但非综述类论文数量多出近六倍。
- 政策可能使计算机与社会领域论文减少约50%,建议基于数据决策。
arXiv最近禁止在计算机科学领域上传未发表的综述类论文,理由是此类论文中存在大量LLM生成内容。然而该决定缺乏量化证据。本文通过两种高质量检测方法,测量了近年来综述与非综述研究论文中LLM生成内容的比例。结果表明,两类论文中LLM生成内容均显著增加,且综述类论文中占比更高。但从生成论文数量看,非综述类论文中的估计值几乎为综述类的六倍。此外,该政策将对特定领域造成更大影响,如计算机与社会子领域论文量或下降约50%。本研究提供了一个基于证据的评估框架,并公开代码以支持未来研究:https://github.com/yanaiela/llm-review-arxiv。
原文摘要 · Abstract (English)
ArXiv recently prohibited the upload of unpublished review papers to its servers in the Computer Science domain, citing a high prevalence of LLM-generated content in these categories. However, this decision was not accompanied by quantitative evidence. In this work, we investigate this claim by measuring the proportion of LLM-generated content in review vs. non-review research papers in recent years. Using two high-quality detection methods, we find a substantial increase in LLM-generated content across both review and non-review papers, with a higher prevalence in review papers. However, when considering the number of LLM-generated papers published in each category, the estimates of non-review LLM-generated papers are almost six times higher. Furthermore, we find that this policy will affect papers in certain domains far more than others, with the CS subdiscipline Computers & Society potentially facing cuts of 50%. Our analysis provides an evidence-based framework for evaluating such policy decisions, and we release our code to facilitate future investigations at: https://github.com/yanaiela/llm-review-arxiv.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。