不用深度学习也能高精度识别ChatGPT改写的新闻文章。
The power of text similarity in identifying AI-LLM paraphrased documents: The case of BBC news articles and ChatGPT
- 基于文本模式相似性设计检测算法,不依赖深度学习。
- 在2224篇真实新闻与同量级生成文上达到96.2%以上准确率。
- 可定位侵权来源为ChatGPT,适合版权保护场景使用。
生成式AI改写文本可能用于盗版侵权,剥夺原创内容创作者的收益。尽管此类恶意应用日益增多,相关学术研究仍不足。本文展示了一种基于模式相似性的检测方法,不仅能识别文章是否为AI改写,更可确定侵权源为ChatGPT。该方法在自建基准数据集上测试,包含来自BBC的2224篇真实新闻及同数量由ChatGPT生成的改写文章,覆盖五个新闻类别。结果显示,该无需深度学习的模式相似性方法,在准确率、精确率、敏感度、特异度和F1值上均达96.23%以上。
原文摘要 · Abstract (English)
Generative AI paraphrased text can be used for copyright infringement and the AI paraphrased content can deprive substantial revenue from original content creators. Despite this recent surge of malicious use of generative AI, there are few academic publications that research this threat. In this article, we demonstrate the ability of pattern-based similarity detection for AI paraphrased news recognition. We propose an algorithmic scheme, which is not limited to detect whether an article is an AI paraphrase, but, more importantly, to identify that the source of infringement is the ChatGPT. The proposed method is tested with a benchmark dataset specifically created for this task that incorporates real articles from BBC, incorporating a total of 2,224 articles across five different news categories, as well as 2,224 paraphrased articles created with ChatGPT. Results show that our pattern similarity-based method, that makes no use of deep learning, can detect ChatGPT assisted paraphrased articles at percentages 96.23% for accuracy, 96.25% for precision, 96.21% for sensitivity, 96.25% for specificity and 96.23% for F1 score.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。